0% found this document useful (0 votes)
5 views17 pages

Biochemistry Unit3 Part2

The document discusses methods for breaking disulfide bonds in proteins and the subsequent sequencing of peptides. It details the use of various proteases and reagents to cleave peptide bonds and the techniques for sequencing the resulting fragments, including mass spectrometry. Additionally, it highlights the advancements in DNA sequencing that allow for the deducing of amino acid sequences from genetic information.

Uploaded by

sila.atayildiz
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views17 pages

Biochemistry Unit3 Part2

The document discusses methods for breaking disulfide bonds in proteins and the subsequent sequencing of peptides. It details the use of various proteases and reagents to cleave peptide bonds and the techniques for sequencing the resulting fragments, including mass spectrometry. Additionally, it highlights the advancements in DNA sequencing that allow for the deducing of amino acid sequences from genetic information.

Uploaded by

sila.atayildiz
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

[ 96 ] Amino Acids, Peptides, and Proteins

Disulfide bond _J
1 (cystine) ..J
7
TH2SH NH O=C
1 1
CHOH HC-CH2-S-S-CH2-CH
1 1 . 1
CHOH C=O HN
1
CH2SH f 7ı
Dithlothreitol (D'IT)

\_ r1 f
NH O O O=C ~NH O=C
1 11 1 1
HC-CH2-s-o- -o-i-cH2-6H HC-CH2-SH HS-CH2-CH
l 11 il 1 1 1
C=O O O HN C=O HN~
,J Cysteic acid
7ı (
ı
residues carboxyınethylation
by

FIGURE 3-26 Breaking disulfıde bonds in proteins. Two com-


iodoacetate
_;
~NH O=C
mon methods are illustrated. Oxidation of a cystine residue 1
with performic acid produces two cysteic acid residues. Re- H6-CH2-S-CH2-coo- -ooc-CH2-S-CH2-CH
1
duction by dithiothreitol or (3-mercaptoethanol to form Cys l
C=O HN~
residues must be followed by further modification of the reac-
tive -SH groups to prevent re-formation of the disulfide bond.
( Carboxymethylated
cysteine
Acetylation by iodoacetate serves this purpose. residues

TABLE 3-7
Reagent (biological source)* Cleavage pointi
Trypsin (bovine pancreas) Lys, Arg (C)
SubmaxiUarus protease (mouse submaxillary gland) Arg (C)
Chymotrypsin (bovine pancreas) Phe, Trp, Tyr (C)
St,a,phylococcus aureus V8 protease (bacterium S. aureus) Asp, Glu (C)
Asp-N-protease (bacterium Pseudomonas fragi) Asp, Glu (N)
Pepsin (porcine stomach) Leu, Phe, Trp, Tyr (N)
Endoproteinase Lys C (bacteriuı:n Lysobacter erızymogenes) Lys (C)
Cyanogen bromide Met (C)

•Ali reagents except cyanogen bromide are proteases. Ali are available from commerclal sources.
tResidues furnishing the primary recognition point for the protease or reagent; peptide bond cleavage occurs on either the car-
bonyt (C) or the amino (N) side of the indicated amino acid residues,

Arnong proteases, the ctigestive enzyme trypsin cat- terminal Lys or Arg. The fragments produced by trypsin
alyzes the hydrolysis of only those peptide bonds in (or other enzyme or chernical) action are then sepa-
which the carbonyl group is contributed by either a Lys rated by chromatographic or electrophoretic methods.
oran Arg residue, regardless of the length or amino acid
sequence of the chain. The nurnber of smaller peptides Seqnencing the Peptides Each peptide fragrnent re-
produced by trypsin cleavage can thus be precticted sulting from the action of trypsin is sequenced separately
frorn the total nurnber of Lys or Arg residues in the by the Edman procedure.
original polypeptide, as determined by hydrolysis of an
intact sample (Fig. 3-27). A polypeptide with three Ordering the Peptide Fragments The order of the
Lys and/or Arg residues (as in Fig. 3-27) wi1l usually "trypsin fragments" in the original polypeptide chain
yield four srnaller peptides on cleavage with trypsin. must now be deterrnined. Another sample of the intact
Moreover, all except one of these wi1l have a carboxyl- polypeptide is cleaved into fragments using a dif[erent
3.4 The Structure of Proteins: Primary Structure [ 97]

enzyme or reagent, one that cleaves peptide bonds at amino acid has been identified before the original cleavage
points other th8:Il those cleaved by trypsin. For exarnple, of the protein, this information can be used to establish
cyanogen bromıde cleaves only those peptide bonds in which fragment is derived from the amino terminus. The
which the carbonyl group is contributed by Met. The two sets of fragments can be compared for possible errors
fragments resulting from this second procedure are in deterrnining the amino acid sequence of each fragment.
then separated and sequenced as before. If the second cleavage procedure fails to establish continu-
The amino acid sequences of each fragment obtained ity between all peptides from the first cleavage, a third or
by the two cleavage procedures are examined, with the even a fourth cleavage method must be used to obtain a set
objective of finding peptides from the second procedure of peptides that can provide the necessaıy overlap(s).
whose sequences establish continuity, because of overlaps,
between the fragments obtained by the first cleavage pro- Locating Disulfide Bonds If the primary structure
cedure (Fig. 3-27). Overlapping peptides obtained from the includes disulfıde bonds, their locations are determined
second fragmentation yield the correct order of the peptide in an additional step after sequencing is completed. A
fragments produced in the first. If the amino-terminal sarnple of the protein is again cleaved with a reagent

s-S Procedure Result Conclusion


ı+1/'.~ hydrolyze; separate R 1
~
A 5 H 2 Polypeptide has 38
-~ amino acids
C 2 I 3 S 2 amino acid residues. Tryp-
D 4 K 2 T 1 sin will cleave three times
E 2 L 2 V 1 (at one R (Arg) and two
F 1 M 2 y 2 K (Lys)) to give four frag-
Polypeptide G 3 p 3 ments. Cyanogen bromide
will cleave at two
M (Met) to give three
fragments.

react with FDNB; hydrolyze;


separate amino acids
2,4-Dinitrophenylglutamate E (Glu) is amino-
reduce
detected terminal residue.
disıılfide
bonds (if present)

cleave with trypsin;


@ GASMALfK @ placed at amine
separate fragments; sequenca
by Edman degradation
§ EGAAYHDFEPIDPR
terminus because it
begins with E (Glu).

§ DCVHSD § placed at carboxyl

8 YLIACGPMTK
terminus because it
does not end with
R (Arg) or K (Lys).

cleave with cyanogen


bromide; separate fragments;
@ EGAAYHDFEPTT)PRGASM @ overlaps with
sequence by Edman degradation
§ TKDCVHSD § and§ , allowing
them to be ordered.
§ ALIKYLIACGPM

eatablish
sequence
Amino
s Carboxyl
EGAAYHDFEPIDPRGASMALIKYLIACGPMTKDCVHSD terminus
terminus

FIGURE 3-27 Cleaving proteins and sequencing and ordering the thus only one possibility for location of the disulfide bond. in polypep-
peptide fragments. First, the amino acid composition and amino- tides with three or more Cys residues, the position of disulfide bonds
terminal residue of an intact sample are determined. Then any disulfide can be determined as described in the text. (The one-letter symbols for
bonds are broken before fragmenting so that sequencing can proceed amino acids are given in Table 3-1.)
efficiently. in this example, there are only two Cys (C) residues and
[98] Amlno Acids, P,ptid,s, and Proteins

Arnino acid Gln-Tyr-Pro-Thr-Ile-Trp


sııchas trypsin, this time \\ithout first breaking tlıe sequence (protein) .-,r--,r-ır-ı.--ı
disulfide bonds. The resulting peptides are separated by DNA sequence (gene) ~TATCCTACGA'f'ITGG
clectrophoresis and compared with tlıe original set of sequences.
ONA and amino acid
peptides generated by trypsin. For eaclı disulfide bond, FIGURE 3-28 Correspondence of .
. . . oded b specifıc sequence of three nucleotıdes
two of the original peptides will be missing and a new, Each amıno acıd ıs ene ya . . . . h ter 27
largcr peptide \\ill appear. The two missing peptides in DNA. The genetic code is descrıbed ın detaıl ın C ap .
represent the regions of tlıe intact polypeptide tlıat are cı the
. available, sequencin"''O DNA can be
linked by the disulfide bond. When t1ıe gene ıs . .

faster and more accurate than sequen :~ ~he proteın.


Amino Add Sequences Can Also Be Deduced by Most proteins are now sequenced in tlıis ın~ect way. lf
Other Methods tlıe gene has not been isolated, direct se~uenc~ of pep-
tides is necessary, and tlıis can provide informail~tiobnl \the
The approach outlined above is not the only way to Iocation of disulfide bonds, for example) not av ~ e ın_ a
determine amino acid sequences. New metlıods based on DNA sequence. in addition, a knowle~e of the amıno ac~d
mass spectrometry permit the sequencing of short n a part ofa polypept ide can greatly facil-
sequence of eve .
polypeptides (20 to 30 amino acid residues) injust a few "tate tlıe isolation of the corresponding gene (Chapter 9).
1
minutes (Box 3-2). in addition, witlı tlıe development of The array of methods now availa~le to anal~e- b~th
rapid DNA sequencing metlıods (Chapter 8), tlıe elucida- proteins and nucleic acids is ushering ın a new discıpline
tion of the genetic code (Chapter 27), and tlıe develop- of "whole cell biochemistry." The complete sequence of
ment of techniques for isolating genes (Chapter 9), an organism's DNA, its genome, is no~ availabl~ for or-
researchers can deduce the sequence ofa polypeptide by ganisms ranging from viruses to bactena to multıcellular
determinirıg the sequence of nucleotides in tlıe gene tlıat eukaryotes (see Table 1-2). New genes are being discov-
codes for it (Fig. 3-28). The techniques used to deter- ered by the thousands, including many that encode
mine protein and DNA sequences are complementarY. (continu ed on page 100)

- METHODS Nff fiiM a+ -M 4A Mll 'it


The mass spectrometer has long been an indispensable
tool in chemistry. Molecules to be analyzed, referred to as
analytes, are .first ionized in a vacuum. When tlıe newly
ing multiply charged macromolecular ions are thus
introduced nondestructively into the gas phase. This
technique is called electrospray ionization mass
spectrometry, or ESi MS. Protons added during pas-
charged molecules are introduced into an electric and/or sage through the needle give additional charge to the
magnetic field, their paths through the [Link] are a function macromolecule. The mlz of the molecule can be ana-
of their mass-to-charge ratio, mlz. This measured prop-
lyzed in the vacuum chamber.
erty of the ionized species can be used to deduce the Mass spectrometıy provides a wealth of information
mass (M) of the analyte with veıy high precision.
for proteomics research, enzymology, and protein chem-
Although mass spectrometıy has been in use for
istry in general. The techniques require only miniscul e
many years, it could not be applied to macromolecules
amounts of sample, so they can be readily applied to the
such as proteins and nucleic acids. The mlz measure-
small amounts of protein that can be extracted from a
ments are made on molecules in the gas phase, and the
two-dimensional electrophoretic gel. The accurately
heating or other treatment needed to transfer a macro-
measured molecular mass ofa protein is one of the crit-
molecule to the gas phase usually caused its rapid de-
composition. In 1988, two different techniques were ical parameters in its [Link]. ünce the mass of a
developed to overcome this problem. In one, proteins protein is accurately known, mass spectrometry is a
are placed in a light-absorbing matrix. With a short pulse convenient and accurate method for detecting changes
of laser light, the proteins are ionized and then desorbed in mass due to the presence of bound cofactors, bound
from the matrix into the vacuum system. This process, metal ions, covalent modifıcations, and so on.
known as matrix-assisted laser desorption/ioniza- The process for determining the molecular mass of
tion mass spectrometry, or MALDI MS, has been a protein with ESI MS is illustrated in Figure 1. As it is
successfully used to measure the mass ofa wide range of injected into the gas phase, a protein acquires a variable
macromolecules . in a second and equally successful number of protons, and thus positive charges, from the
metlıod, macromolecules in solution are forced directly solvent. This creates a spectrwn of species with differ-
from the liquid to gas phase. A solution of analytes is ent mass-to-charge ratios. Each successive peak corre-
passed through a charged needle that is kept at a high sponds to a species that differs from that of its
electrical potential, dispersing the solution into a [Link] neighboring peak by a charge difference of 1 and a mass
mist of charged microdroplets. The solvent surrounding difference of 1 (1 proton). The mass of the protein can
the macromolecules rapidly evaporates, and the result- be determined from any two neighboring peaks.
5
3.4 The Structure of Proteins: Primary Structure [ 99]

The measured mlz of one peak is


Mass
Glass Sample spectrometer

where M is the mass of the protein, n2 is the number of


capillary solution
1
charges, and X is the mass of the added groups (protons
in this case). Similarly for the neighboring peak,
High
(m/z)ı = M + (n 2 + l)X voltage
n2 +1
Vacuum
We now have two unknowns (M and n 2) and two equa- (a) interface
tions. We can solve first for ¾ and then for M:
(m/z)2 -X
n2= - -....::.__ _
(m/z)2 - (m/z)ı
,~

·:LL
M = n 2 [(mlzh - X]

This calculation using the mlz values for any two peaks 50+
in a spectrum such as that shown in Figure ı b usually l 100 !
provides the mass of the protein (in this case, aerolysin ~
k; 47,342 Da) with an error of only ±0.01%. Generating -~ 75 ,... 40+ 47,000 48,000
.s ! Mr
several sets of peaks, repeating the calculation, and av- .s
CI)
eraging the results generally provides an even more ac- :oaı> 50 ,... 30+
curate value for M. Computer algorithms can transforrn !
the mlz spectrum into a single peak that also provides a ~
25 ,...
very accurate mass measurement (Fig. lb, inset).
Mass spectrometry can also be used to sequence
short stretches of polypeptide, an application that has o
.,,til.
1
w
1 1 1
J..1 1
J ıt
emerged as an invaluable tool for quickly identifying un- 800 1,000 1,200 1,400 1,600
lrnown proteins. Sequence inforrnation is extracted using (b) mlz
a technique called tandem MS, or MS/MS. A solution FIGURE 1 Electrospray mass spectrometry of a protein. (a) A protein
containing the protein under investigation is first treated solution is dispersed into highly charged droplets by passage through a
with a protease or chemical reagent to hydrolyze it to a needle under the influence ofa high-voltage electric field. The droplets
mixture of shorter peptides. The mixture is then injected evaporate, and the ions (with added protons in this case) enter the
into a device that is essentially two mass spectrometers mass spectrometer for m/z measurement. The spectrum generated (b)
in tandem (Fig. 2a, top). In the first, the peptide mixture is a family of peaks, with each successive peak (from right to left) cor-
is sorted and the ionized fragments are manipulated so responding to a charged species increased by 1 in both mass and
that only one of the several types of peptides produced charge. lnset: a computer-generated transformation of this spectrum.
by cleavage emerges at the other end. The sample of the
selected peptide, each molecule of which has a charge
somewhere along its length, then travels through a vac- The second mass spectrometer then measures the
uum chamber between the two mass spectrometers. In mlz ratios of ali the charged fragments (uncharged
this collision celi, the peptide is further fragmented by fragments are not detected). This generates one or
high-energy impact with a "collision gas," a small amount more sets of peaks. A given set of peaks (Fig. 2b) con-
of a noble gas such as helium or argon that is bled into sists of ali the charged fragments that were generated
the vacuum chamber. This procedure is designed to frag- by breaking the same type of bond (but at different
ment many of the peptide rnolecules in the sample, with points in the peptide) and are derived frorn the same
each individual peptide broken in only one place, on av- side of the bond breakage, either the carboxyl- or
erage. Most breaks occur at peptide bonds. This frag- amino-terminal side. Each successive peak in a given set
mentation does not involve the addition of water (it is has one less amino acid than the peak before. The dif-
done in a near-vacuum), so the products rnay include ference in mass from peak to peak identifies the amino
rnolecular ion radicals such as carbonyl radicals (Fig. 2a, acid that was lost in each case, thus revealing the se-
bottorn). The charge on the original peptide is retained quence of the peptide. The only ambiguities involve
on one of the fragments generated from it. leucine and isoleucine, which have the same mass.
(conti nued on next page)
[1 oo] Amino Acids, Peptides, and Proteins

- METHODS

MS-1
Collision
celi MS-2 Detector
FIGURE 2 Obtaining protein sequence information with tandem MS.
(a) After proteolytic hydrolysis, a protein solution is injected into a mass
spectrometer (MS-1 ). The different peptides are sorted so that only one
~ ~ @) type is selected for further analysis. The selected peptide is further frag-
Electrospray Separation Breakage
;, ru,atioo, 1 mented in a chamber between the two mass spectrometers, and m/z for
each fragment is measured in the second mass spectrometer (MS-2).
Many of the ions generated during this second fragmentation result
from breakage of the peptide bond, as shown. These are called b-type
or y-type ions, depending on whether the charge is retained on the
amino- or carboxyl-terminal side, respectively. (b) A typical spectrum
R1 O R8 O R6 with peaks representing the peptide fragments generated from a sample
H H I b
1 il il ~ H H 1 ~O
11,N-C-C-N -C-C-N- C-C N-C-C-N- C-C of one small peptide (1 Oresidues). The labeled peaks are y-type ions.
H H I il H H I il H 'o-
R2 O ~ O The large peak next to y5" is a doubly charged ion and is not part of the
y
y set. The successive peaks differ by the mass ofa particular amino acid
in the original peptide. in this case, the deduced sequence was
R1 O R8 O RI
1 il H H I il H H 1 ~O Phe-Pro-Gly-Gln--(lle/Leu)-Asn-Ala-Asp-(lle/Leu)-Arg. Note the am-
RıN-C-C-N-C-C-N-C-C• •N-C-C-N-C -C
H H I il H H I il H 'o- biguity about ile and Leu residues, because they have the same molec-
Rı O R4 O ular mass. in this example, the set of peaks derived from y-type ions
(a) predominates, and the spectrum is greatly simplified as a result. This is
because an Arg residue occurs at the carboxyl terminus of the peptide,
and most of the positive charges are retained on this residue.

100 . ..
ı
Y2" fragments can be unambiguously distinguished from that
~
l
75 ·I
consisting of the arnino-ternıinal fragments. Because th~
b
·; Ya" l bond breaks generated between the spectrometers (in the

ı::
.s 50 collision cell) do not yield full [Link] and amino groups
.s y/ Ys" .j at the sites of the breaks, the only intact a-amino arıd
CI) 11·
.::... 25
Ys°
j a-carboxyl groups on the peptide fragments are those at
<il

~ 1 the very ends (Fig. 2a). The two sets of fragments ca.11
o thereby be identified by the resulting slight differences in
200 400 600
(b) mlz mass. The amino acid sequence derived from one set can
be confirmed by the other, improving the confidence in the
sequence information obtained.
Even a short sequence is often enough to pennit un-
The charge on the peptide can be retained on either ambiguous association of a protein with its gene, if the
the [Link]- or arnino-ternıinal fragment, and bonds other gene sequence is known. Sequencing by mass spectrom-
tlıan the peptide bond can be broken in the [Link] etry cannot replace the Edman degradation procedure
process, with the result that multiple sets of peaks are usu- for the sequencing of long polypeptides, but it is ideal
alJy generated. The two most prominent sets generally for proteornics research aimed at cataloging the hun-
consist of charged fragments derived from breakage of the dreds of cellular proteins that rnight be separated on a
peptide bonds. The set consisting of the [Link]-ternıinal two-dimensional gel.

(continuedfrom page 98)


proteins with no known function. To describe the entire importance. There are three ways to obtain a peptide: (1)
protein complement encoded by an organism's DNA, purification from tissue, a task often made difficult by the
researchers have coined the term proteome. As de- vanishingly low concentrations of some peptides; (2) ge-
scribed in Chapter 9, the new disciplines of genornics and netic engineering (Chapter 9); or (3) direct chernical syn-
proteornics are complementing work carried out on cellu- thesis. Powerful techniques now make direct cheınical
lar intermediary [Link] and nucleic acid metabolism synthesis an attractive option in many cases. In addition to
to provide a new and increasingly complete picture of commercial applications, the synthesis of specific peptide
biochernistry at the Ievel of cells and even organisms. portions of larger proteins is an increasingly important tool
for the study of protein structure and function.
Small Peptidesand ProteinsCanBe Chemically Synthesized The complexity of proteins makes the traditional
Many peptides are potentially useful as pharmacologic synthetic approaches of organic chernistry impractical
agents, and their production is of considerable commercial for peptides with more than four or five arnino acid 1
3.4 The Structure of Proteins: Prlmary Structure [ 101]

residues. üne problem is u,e difficulty of purifying the 1


product [Link] each step. that used for chromatographic procedures. The peptid e
is built up on this support one amino acid at a time,
The major breakthrough in this technology was pro-
through a standard set of reactions in a repeating cycle
vided by R. Bnıce Merrifıeld in 1962. His innovation in-
(Fig. 3-29). At each successive step in the cycle, pro-
volved synthesizing a peptide while keeping it attached
tective chemical groups block unwanted reactions.
at one end to a solid support. The support is an insoluble
The technology for chemical peptide synthesis is now
polymer (resin) contained within a column, similar to
automated. As in the sequencing reactions considered

FIGURE 3-29 Chemical synthesis ofa peptide on an insoluble polymer support


.
Reactions (D through ©are necessary for the formation of each
0 R1 0 peptide bond.
l The 9-fluorenylmethoxycarbonyl (Fmoc) group (shaded blue)
il I il prevents un-
CH2 -0-C :-N- CH- c-o- wanted reactions at the a-amino group of the residue (shaded
red). Chemical
l synthesis proceeds from the carboxyl terminus to the amino terminu
s, the re-
.J H verse of the direction of protein synthesis in vivo (Chapter 27).

Amino acid
Fmoc residue

-t
Insoluble

Amino acid 1 with


a-amino group protected
1
Cl-CH2
. -0-0 polystyrene
bead

Fmoc ·- R1 il
-N- CH- O
_ Attachment of carboxyl-terminal
by Fmoc group C- 0 (D amino acid to reactive
,J H
I group on resin.
...1 , ·-- · , cı-

R1 O
ı Fmoc ~ N-6H-~-O-CH2 l\.____r\ ., ___________ _____\
1
_. ,..,- H ______... ~ - 1
R2 O ,._.ı. . r·
ı

1 il
Fmoc - N- CH- c-o- Protecting group is removed
1 ® by fluahing with solution
H containing a mild organic base.

Amino acid 2 with


protected
a-amino group is
R1 o· ·
Dicyclohexylcarbodiiınide
® activated at ' + I il ~
carboxyl group H3N- CH-C -~ C H 2 ~ -
(DCC) byDCC.
1

, " R 2
O
Q NH ©
4
a-Aminogroupofamino
acid 1 attacks activated
carboxyl group of amino acid
1
1
1
1
1
1
1
. 1 1

6
11 1 2 to form peptide bond.
Fmoc -N- CH- C-0- C 1
1
1 il

o~-~-~-o
1
o 1
1
- H 1
1
1
1
H H 1
Dicyclohexylurea byproduct 1
1
1
1
~ Q ~ Ü •~~® to@ 1

~ N- 6H-~ -N-6 H-~- O-CH 2 ~ --~~..:


1
Fmoc '1~~-n.:':!~!.1!_-✓'
1
1 1 """\J_ '-ili
- H H

l-0-0'
Completed peptide is
HF deprotected 88 in
@ reaction®; HF cleaves
ester linkage between
peptide and resin.

R2 O R' O ı ,
+ CH-
H N- 1 11
C-N- 1
CH- cil - o - + F-CH 2
R. Bruce Merrifield 3 1 -
1921-2 006 H
[1 oı] Amino Acids, P,ptides, and Proteins

Exactly how Uıe amino acid sequence de~ermirn~s


TABLE 3-8 three-dimensional structure is not understood ın detail,
nor can we always predict function from sequence. How-
ever protein families that have some shared structural or
Overall yield of final functional features can be readily identifıed on the hasis of
peptlde (%) when the
yield of each step is: amino acid sequence similarities. Individual proteins are
Number of residues in assigned to families based on the degree ?f similarity in
the final polypeptide 96.0% 99.8% amino acid sequence. Members of a family are usually
11 66 98 identical across 25% or more of their sequences, and pro-
teins in these families generally share at least some stnıc­
21 44 96
tural and functional characteristics. Sorne families are
defıned, however, by identities involvirıg only ~ few amino
31 29 94
51 13 90 acid residues that are critical to a certain function. A num-
100 1.8 82 ber of similar substructures, or "dornains" (to be defıned
more fully in Chapter 4), occur in rnany fun:tionally unre-
lated proteins. These domains often fold ınto st~~tural
above, the most important limitation of the process is the corıfigurations that have an unusual degree of stability or
effıciency of each chemical cycle, as can be seen by cal- that are specialized for a certain environrnent. Evolution-
culating the overall yields of peptides of various lengths ary relationships can also be inferred ~om ~~ structural
when the yield for addition of each new amino acid is and functional similarities within proteın families.
96.0% versus 99.8% (Table 3-8). Incomplete reaction at Certain amino acid sequences serve as signals that de-
one stage can lead to formation of an impurity (in the termine the celiular location, chernical rnodifıcation, and
form ofa shorter peptide) in the next. The chemistry has half-life of a protein. Special signal sequences, usually at
been optimized to permit the synthesis of proteins of 100 the amino terminus, are used to target certain proteins for
amino acid residues in a few days in reasonable yield. A export frorn the celi; other proteins are targeted for distıi­
very similar approach is used to synthesize nucleic acids bution to the nucleus, the celi surface, the cytosol, or
(see Fig. 8-35). It is worth noting that this technology, other celiular locations. Other sequences act as attach-
impressive as it is, stili pales when compared with biolog- ment sites for prosthetic groups, such as sugar groups in
ical processes. The same 100-residue protein would be glycoproteins and lipids in lipoproteins. Sorne of these sig-
synthesized with exquisite fıdelity in about 5 seconds in a nals are well characterized and are easily recognized in the
bacterial celi. sequence ofa newly characterized protein (Chapter 27).
A variety of new methods for the effıcient ligation
Qoining together) of peptides has made possible the as-
KEY CONVENTION: Much of the functional information
sembly of synthetic peptides into larger polypeptides and
encapsulated in protein sequences comes in the form of
proteins. With these rnethods, novel forms of proteins can
consensus seqnences. This terrn is applied to such
be created with precisely positioned chemical groups, in-
sequences in DNA, RNA, or protein. When a series of
cluding those that might not normally be found in a celiu-
related nucleic acid or protein sequences are cornparec\,
lar protein. These novel forms provide new ways to test
a consensus sequence is the one that reflects the most
theories of enzyrne catalysis, to create proteins with new
cornrnon base or amino acid at each position . Parts of
chernical properties, and to design protein sequences
the sequence that have particularly good agreement often
that will fold into particular structures. This last applica-
represent evolutionarily conserved functional domains . A
tion provides the ultimate test of our increasing ability
range of mathernatical tools available on the Intemet can
to relate the primary structure ofa peptide to the three-
be used to generate consensus sequences, or identify
dirnensional structure that it takes up in solution.
them in sequence databases. Box 3-3 illustrates cornrnon
conventions for displaying consensus sequences. ■
AminoAcidSequences Provide lmportant
Biochemical lnformation
Protein Sequences Can Elucidate the History
Knowledge of the sequence of amino acids in a protein of Life on Earth
can offer insights into its three-dirnensional structure
and its function, cellular location, and evolution. Most of The simple string of letters denoting the arnino acid se-
these insights are derived by searching for similarities quence ofa protein belies the wealth of inforrnation this
between a protein of interest and previously studied sequence holds. As rnore protein sequences have be-
proteins . Thousands of sequences are known and avail- corne available, the developrnent of rnore powerful
able in databases accessible through the Internet. A methods for extracting information frorn thern has be-
cornparison ofa newly obtained sequence with this large corne a major biochemical enterprise. Analysis of the
bank of stored sequences often reveals relationships information available in the rnany, ever-expanding bio-
both surprising and enlightening. logical databases, including gene and protein sequences
[103]

-
3.4 The Structure of Proteins: Primary Structure

Consensus sequences can b


ways. To illustrate two types ~ represented in several
0 of an element of the pattern is indicated by foll .
examples of consensus sequen~~:v:~tions'. we ~se two that element with a number or range of numbe~~
(a) an ATP binding ' own ın Fıgure 1·
12-2) and (b)
a c}~rub~turdingecalled a p loop (see Bo~
tween parentheses. In (a), for example x( )
x d (2 4) 4 means x-x-
-x, an · x , . means x-x, or x-x-x, or, x-x-x-x
- ın structure call d E Wh
hand (see Fig. 12-11) The r . e an F Pattem ıs restncted to either the arrun· or b. ena
adapted from those us~d by t~es descrıbed here are • 0 car oxy1ter-
mınus ofa sequence, that pattern starts wı·th <
·t PRO sequence comparison ·h • d
websı e or en s
SITE ([Link]/prosite)· wıt . >, respectıvely (not so for either example here). A
standard one-letter codes for the anun· , ~hdey use the perıod ends the pattern. Applying these rules to th
o acı s.
[AG)-x(4)-G-K-[ST]. consensus s:~uence in (a), either A or G can be found a~
the first -~osıtıon. Any amino acid can occupy the next

~ :b
Ol
=-Cli
234567s
four posıtıons, followed by an invariant G and an invari-
ant K. The last position is either S or T.
S~quence logos provide a more infonnative and
(a) N C grap~c representation of an amino acid (or nucleic acid)
D-{W}-IDNSJ-{ILVFYW}-IDENSTGJ-IDNQGHRKJ-{ GP}- multıple sequence alignment. Each logo consists of a

4--------~=~-~
3
[LIVMCJ-IDENQSTAGCJ-x(2)-IDEI- ILIVMFYwJ. stack of symbols for each position in the sequence. The
overall height of the stack (in bits) indicates the degree
of sequence conservation at that position while the
-~ 2 height of each symbol in the stack indicates the relative
ıı:ı
1 fr:quency of that amino acid (or nucleotide). For amino
o 1 acıd sequences, the colors denote the characteristics of
2 3 4 5 6 7 8 9 10 11 12 13 the amino acid: polar (G, S, T, Y, C, Q, N) green; basic (K,
(b) N
C R, H) blue; acidic (D, E) red; and hydrophobic (A V L I
FIGURE 1 Representations of two consensus sequences. (a) p loop, an P, W, F, M) black. The classifıcation of amino acid~ ~ t~
ATP-binding structure; (b) EF hand, a Caı+-binding structure. scheme is somewhat different from that in Table 3- ı and
Figure 3-5. The amino acids with aromatic side chains
In one type of consensus sequence designation
are subsumed into the nonpolar (F, W) and polar (Y)
(shown at the top of (a) and (b)), each position is sepa-
classifıcations. Glycine, always hard to group, is assigned
rated from its neighbor by a hyphen. A position where
to the polar group. Note that when multiple amino acids
any amino acid is allowed is designated x. Ambiguities are acceptable at a particular position, they rarely occur
are indicated by listing the acceptable amino acids for a with equal probability. üne or a few usually predominate.
given position between square brackets. For example, in The logo representation makes the predominance clear,
(a} [AG] means Ala or Gly. If all but a few amino acids and a corıserved sequence in a protein is made obvious.
are allowed at one position, the amino acids that are not However, the logo obscures some amino acid residues
allowed are listed between curly brackets. For example, that may be allowed at a position, such as the Cys that
in (b) (W} means any amino acid except Trp. Repetition occasionally occurs at position 8 of the EF hand in (b).

and macromolecular structures, has given rise to the The field of molecular evolution is often traced to
new field of bioinformatics. üne outcome of this disci- Emile Zuckerkandl and Linus Pauling, whose work in
pline is a growing suite of computer programs, many the mid-1960s advanced the use of nucleotide and pro-
readily available on the Internet, that can be used by any tein sequences to explore evolution. The premise is de-
scientist, student, or knowledgeable layperson. Each ceptively straightforward. If two organisms are closely
protein's function relies on its three-dimensional struc- related, the sequences of their genes and proteins
ture, which in turn is determined largely by its primary should be similar. The sequences increasingly diverge as
structure. Thus, the biochemical information conveyed the evolutionary distance between two organisms in-
by a protein sequence is limited only by our own under- creases. The promise of this approach began to be real-
standing of structural and functional principles. The ized in the 1970s, when Carı Woese used ribosomal RNA
constantly evolving tools of bioinformatics make it sequences to define the Archaea as a group of living
possible to [Link] functional segments in new_proteins organisms distinct from the Bacteria and Eukarya (see
and help establish both their sequence and their struc- Fig. 1- 4). Protein sequences offer an opportunity to
tural relationships to proteins already in the databases. greatly refine the available information. With the advent
On a different level of inquiry, protein sequences are of genome projects investigating organisms from bacte-
beginning to tell us how the proteins evolved and, ulti- ria to humans, the number of available sequences is
mately, how life evolved on this planet. growing at an enormous rate. This information can be

[ 104] Amino Acids, Peptides, and Proteins

used. to trace biological history. The challenge is in are called orthologs. The process of tracing evolution
leanung to read the genetic hieroglyphics. involves fırst identifying suitable families of homolo-
~~olution has not taken a sirnple linear path. Corn- gous proteins and then using thern to reconstruct evo-
plexıties abound in any atternpt to mine the evolution- lutionary paths.
aıy ~orrnation stored in protein sequences. Fora given Hornologs are identifıed through the use of increas-
proteın, the amino acid residues essential for the activ- ingly powerful computer prograrns that can directıy
ity of the protein are conserved over evolutionary time. cornpare two or rnore chosen protein sequences, or can
The residues that are less important to function rnay search vast databases to .find the evolutionary relatives
vary over tirne-that is, one arnino acid rnay substitute of one selected protein sequence. The electronic search
for another-and these variable residues can provide process can be thought of as sliding one sequence past
the inforrnation to trace evolution. Arnino acid substitu- the other until a section with a good rnatch is found.
tions are not always randorn, however. At sorne posi- Within this sequence aligrunent, a positive score is as-
tions in the prirnaıy structure, the need to rnaintain signed for each position where the arnino acid residues
protein function rnay rnean that only particular arnino in the two sequences are identical-the value of the
acid substitutions can be tolerated. Some proteins have score varying frorn one program to the next-to provide
more variable amino acid residues than others. For a measure of the quality of the alignrnent. The process
these and other reasons, different proteins can evolve at has some cornplications. Sornetimes the proteins being
düferent rates. compared rnatch well at, say, two sequence segments,
Another complicating factor in tracing evolutionary and these segments are connected by less related se-
history is the rare transfer of a gene or group of genes quences of düferent lengths. Thus the two rnatching
from one organism to another, a process called lateral segments cannot be aligned at the sarne time. To handle
gene transfer. The transferred genes may be quite this, the computer program introduces "gaps" in one of
similar to the genes they were derived frorn in the origi- the sequences to bring the rnatching segments into
nal organism, whereas rnost other genes in the same two register (Fig. 3-30). Of course, ifa suffıcient nurnber of
organisms may be quite distantly related. An example of gaps are introduced, alrnost any two sequences could be
lateral gene transfer is the recent rapid spread of brought into sorne sort of aligrunent. To avoid uninfor-
antibiotic-resistance genes in bacterial populations. The rnative alignments, the prograrns include penalties for
proteins derived from these transferred genes would not each gap introduced, thus lowering the overall align-
be good candidates for the study of bacterial evolution, rnent score. With electronic trial and error, the program
because they share only a very limited evolutionary his- selects the alignment with the optimal score that rnaxi-
tory with their "host" organisms. rnizes identical amino acid residues while minirnizing
The study of molecular evolution generally focuses the introduction of gaps.
on families of closely related proteins. In most cases, Identical arnino acids are often inadequate to identify
the families chosen for analysis have essential func- related proteins or, rnore importantly, to deterrnine how
tions in cellular metabolism that must have been pres- closely related the proteins are on an evolutionary tµne
ent in the earliest viable cells, thus greatly reducing scale. Arnore useful analysis includes a consideration of
the chance that they were introduced relatively re- the chernical properties of substituted arnino acids. When
cently by lateral gene transfer. For example, a protein arnino acid substitutions are found within a protein fam-
called EF-la (elongation factor la) is involved in the ily, rnany of the düferences rnay be conservative-that
synthesis of proteins in ali eukaryotes. A similar pro- is, an arnino acid residue is replaced by a residue havmg
tein, EF-Tu, with the sarne function, is found in bacte- similar chernical properties. For exarnple, a Glu residue
ria. Similarities in sequence and function indicate that may substitute in one farnily rnernber for the Asp residue
EF-la and EF-Tu are members ofa family of proteins found in another; both arnino acids are negatively
that share a common ancestor. The members of protein charged. Such a conservative substitution should logi-
families are called homologous proteins, or ho- cally garner a higher score in a sequence alignrnent than
mologs. The concept of a homolog can be further re- does a nonconservative substitution, such as the replace-
fined. If two proteins in a family (that is, two ment of the Asp residue with a hydrophobic Phe residue.
homologs) are present in the same species, they are re- For most efforts to find homologies and explore evo-
ferred to as paralogs. Homologs from düferent species lutionary relationships, protein sequences (derived either

E. coli TGNR T I AV YDLGGGTFD I s I I E I DEVDGEK TFEV LATNGD TH LGG EDFDS RL I HY L


B. subtilis DEDQ T I LL YDLGGGTF DVS ILE LGDG TFEV RS TAGD NR LGGDDFDQV II DH L
L__J
Gap

FIGURE 3-30 Aligning protein sequences with the use of gaps. Shown bacterial species, E. co/i and Bacillus subtilis. lntroduction ofa gap in ıhe
here is the sequence alignment of a short section of the Hsp70 proteins B. subtilis sequence allows a better alignment of amino acid residues on
(a widespread class of protein-folding chaperones) from two well-studied either side of the gap. ldentical amino acid residues are shaded.
3.4 The Structure of Proteins: Primary Structure [1 os]
Signature sequence
Archaea Halobacterium halobium IGHVD HGK S TMVGR LYET GSVPEHVIEQH
Sulfolobus solfataricus IGHVaHUK S TLVGR L LMDRGFIDEKT KEA
Eukaryotes { Saccharomyces cerevisiae IGHVDSGK S~ TTGH L IYKC GGIDKRTIEKF
.. . Homo sapiens IGHVD S G~S TTTGH L IYKC GGIDKRTIEKF
Gram-posıtive bactenum Bacillus subtilis IGHVD HGK STMVGR ITTV
Gram-negative bacterium Escherichia coli IGHVD HGKTTLTAA ITTV

FIGURE3-31 A signature sequence in the EF-1a/EF-Tu protein family. although the sequences of the insertions are quite distinct for the two
The signature sequence (boxed) is a 12-residue insertion near ıhe groups. The variation in the signature sequence reflects the significant
amino terminus of the sequence. Residues that align in all species are evolutionary divergence that has occurred at this site since it first ap-
shaded yellow. Both archaea and eukaryotes have the signature, peared in a common ancestor of both groups.

directly from protein sequencing or from the sequencing ture sequences for the group in which they are found.
of the DNA encoding the protein) are superior to non- An example of a signature sequence is an insertion of
genic nucleic acid sequences (those that do not encode a 12 amino acids near the amino terminus of the EF-
protein or functional RNA). For a nucleic acid with its la/EF-Tu proteins in all archaea and eukaryotes but not
four different types of residues, random alignment' of non- in bacteria (Fig. 3-31). This particular signature is one
homologous sequences will generally yield matches for at of many biochemical clues that can help establish the
least 25% of the positions. Introduction of a few gaps can evolutionary relatedness of eukaryotes and archaea.
often increase the fraction of matched residues to 40% or Other signature sequences allow the establishment of
more, and the probability of chance alignment of unre- evolutionary relationships among groups of organisms
lated sequences becomes quite high. The 20 different at many different taxonomic levels.
amino acid residues in proteins greatly lower the probabil- By considering the entire sequence of a protein,
ity of uninformative chance alignrnents of this type. researchers can now construct more elaborate evolu-
The programs used to generate a sequence align- tionary trees with many species in each taxonomic
ment are complemented by rnethods that test the relia- group. Figure 3-32 presents one such tree for bacte-
bility ofthe alignments. A cornmon computerized test is ria, based on sequence divergence in the protein
to shuf:fle the amino acid sequence of one of the pro- GroEL (a protein present in ali bacteria that assists in
teins being compared to produce a random sequence, the proper folding of proteins). The tree can be re-
then to instruct the program to align the shuf:fled se- fined by basing it on the sequences of rnultiple pro-
quence with the other, unshuf:fled one. Scores are as- teins and by supplementing the sequence information
signed to the new alignment, and the shuffling and with data on the unique biochemical and physiologi-
aligrunent process is repeated many times. The original cal properties of each species. There are many meth-
aligrunent, before shnffling, should have a score signifi- ods for generating trees, each method with its own
cantly higher than any of those within the distribution advantages and shortcomings, and many ways to rep-
of scores generated by the random alignments; this in- resent the resulting evolutionary relationships. In
creases the confidence that the sequence alignment has Figure 3-32, the free end points of lines are called
identifıed a pair of homologs. Note that the absence of "external nodes"; each represents an extant species,
a significant alignment score does not necessarily mean and each is so labeled. The points where two lines
that no evolutionary relationship exists between two come together, the "internal nodes," represent ex-
prnteins. As we shall see in Chapter 4, three-dimensional tinct ancestor species. in most representations (in-
structural similarities sornetirnes reveal evolutionary cluding Fig. 3-32), the lengths of the lines connecting
relationships where sequence hornology has been the nodes are proportional to the number of amino
wiped away by time. acid substitutions separating one species frorn an-
Use ofa protein family to explore evolution requires other. If we trace two extant species to a comrnon in-
the [Link] of family rnernbers with similar rnolec- ternal node (representing the common ancestor of
ular functions in the widest possible range of organisrns. the two species), the length of the branch connecting
Information frorn the family can then be used to trace each external node to the internal node represents
the evolution of those organisrns. By analyzing the se- the nurnber of amino acid substitutions separating
quence divergence in selected protein families, investi- one extant species from this ancestor. The sum of the
gators can segregate organisms into classes based on lengths of all the line segments that connect an extant
their evolutionary relationships. This inforrnation must species to another extant species through a cornmon
be reconciled with more classical examinations of the ancestor reflects the nurnber of substitutions separat-
physiology and biochemistry of the organisrns. ing the two extant species. To determine how much
Certain segments of a protein sequence may be time was needed for the various species to diverge,
found in the organisrns of one taxonomic group but not the tree must be calibrated by comparing it with in-
in other groups; these segments can be used as signa- formation from the fossil record and other sources.
[106] Amino Acids, Peptides, and Particles

Bacteroides [
[Link]
• [ Chlamydia trachomatis
Chlamydia psittaci
Porp/ıyroınoııa& giııgivalis
Borrelia bu'lldorferi l Spirochaetes
LtptMpira interrogaM

Oj
·c
r
[
Ltgioııellıı pneumophila

Ytnıinia enttrocolitica
Tl,~pJ,llk ...... ,,.,
Baci/luı subtilis
Staphylococcus aureus
Clostridium acetobutylicum
l
Jow
G+C
Oj
·c
...
Q)

"
Oj
.Q
~
Salmonella typhi CIMtridium per{ringen• Q)
Eschtrichia coli
~
Oj
.Q

·~---ı
0 1/J
~
Riı:kdtsia
/3 0
ı:l.

l tıutsugamu,hi

Mycobacterium leprae .
high
G+C ..~
Bradyrhizobiumjaponiı:um Mycobacterium tuberculosı& c:,
a Streptomyce• albu, [genel

Agrobacterium tumefaciena
Zymamanas mabilis

Cyanobacteria and
chloroplasts

0.1 substitutions/site Tritiı:um atstivum ehi.


Brassiı:a napus clıl.
Arabidopsis thaliana clıl.

FIGURE 3-32 Evolutionary tree derived from amino acid sequence vergence observed in the GroEL family of proteins. Also included in this
1
1
comparisons. A bacterial evolutionary tree, based on the sequence di- tree (lower right) are the chloroplasts (ehi.) of some nonbacterial species.
11

,,1 As more sequence information is made available in organism on Earth. The story is a work in progress, of
databases, we can generate evolutionary trees based course (Fig. 3-33). The questions being asked andan-
1 on multiple proteins. And we can refine these trees as swered are fundamental to how humans view themselves
1 additional genomic information emerges from increas- and the world around them. The fıeld of molecular evolu-
'
ingly sophisticated methods of analysis. All of this work tion promises to be among the most vibrant of the scien-
moves us toward the goal of creating a detailed tree of tifıc frontiers in the twenty-fırst century.
life that describes the evolution and relationship of every

LowG+C, Crenarchaeota
gram-positive Thermo- Desulfurococcales
proteales
Thermotogales Sulfolobales
Aquificales ~ Euryarchaeota
Spirochaetes / Halobacteriales
Chlamydiales~ \ ~ Methanosareinales
'• \ ,, Th I t1
De~ococcales .::: - - ~ ' ermop asma a e.s
High G + C,_ ~ - Bacter·a Arch'' Arehaeoglobales
gram-negatıve ı aea
Cyanobacteria ~ -~ , , Methanococcales
. Mıtochon drıa Thermococcales
Proteobactena '-.. Eukarya
. - - - -- Chloroplast~
[Land plantıJ \ ~-,.~- Opisthokonta
Green algae ~ - -

Planta~ı!:oı:~s
Mycetozoans
Pelobionts
1.1/Ğ
',~
'\ ',,~
/ \ ',, ----....::::---Fungi

Ch~:~:=~ates
(multicellular aniınals)
Radiolaria
Entamoebae Cereozoa
Amoebozoa Rhizari
Diplomonads Jalı. b"d Alveolates a
O1
Euglenoids c~to- Stramenopiles
Excavata phytes Haptophytes
Chromalveolata

FIGURE 3-33 Aconsensus tree of life. The tree shown here is based on The tree presents only a fraction of the available information, as well as
analyses of many different protein sequences and additional genomic only a fraction of the issues remaining to be resolved. Each exıanı
features. Branches shown as dashed lines remain under investigation. group shciwn is a complex evolutionary story unıo itself.
Further Reading [101]

SUMMARY 3.4 The Structure of Proteins: consensus sequence 102 homolog 104
bioinformatics 103 paralog 104
Primary Structure lateral gene transfer 104 ortholog 104
■ Differences in protein function result from homologous proteins 104 signature sequence 105
differences in amino acid composition and sequence.
Some variations in sequence are possible for a
particular protein, with little or no effect on function. Further Reading - - - - - - - - -
■ Amino acid sequences are deduced by fragmenting Amino Acids
po]ypeptides into smaller peptides with reagents Dougherty, D.A. (2000) Unnatural arnino acids as probes of protein
known to cleave specific peptide bonds; determining structure and function. Gurr: Opin. Ghem. Biol. 4, 645-652.
the amine acid sequence of each fragment by the Greenstein, J.P. & Wınitz, M. (1961) Ghemistry ojtheAmino
automated Edman degradation procedure; then Acids, 3 Vols, John Wıley & Sons, New York.
ordering the peptide fragments by finding sequence Krell, G. (1997) o-Arnino acids in animal [Link]. Rev.
overlaps between fragments generated by different Biochem. 66, 337-345.
reagents. A protein sequence can also be deduced Details the occurrence of these unusual stereoisomers of
arnino acids.
from the nucleotide sequence of its corresponding
gene in DNA. · Meister, A. (1965) Biochemistry of the Amino Acids, 2nd edn,
Vols 1 and 2, Academic Press, ine., New York.
■ Short proteins and peptides (up to about 100 Encyclopedic treatment of the properties, occurrence, and
residues) can be chemically synthesized. The metabolism of arnino acids.
peptide is built up, one amino acid residue at a time, Peptides and Proteins
while tethered to a solid support.
Creighton, T.E. (1992) Proteins: Structures and Molecular
■ Protein sequences are a rich source of information Properties, 2nd edn, W. H. Freeman and Company, New York.
about protein structure and function, as well as the Very useful general source.
evolution of life on Earth. Sophisticated methods Working with Proteins
are being developed to trace evolution by analyzing Dunn, M.J. & Corbett, J.M. (1996) 'I\vo-dimensional polyacryl-
the resultant slow changes in amine acid sequences amide gel electrophoresis. Methods Enzymol. 271, 177-203.
of hornologous proteins. A detailed description of the technology.
Komberg, A. (1990) Why purify enzyrnes? Methods Enzymol.
182, 1-5.
KeyTerms - - - - - - - - - - - The critical role of classical biochemical methods in a new age.
Turms in bold are defined in the glossary. Scopes, R.K. (1994) Protein Purijication: Principles and
Practice, 3rd edn, Springer-Verlag, New York.
. A good source for more complete descriptions of the principles
amino acids 72 column chromatog-
underlying chromatography and other rnethods.
Rgroup 72 raphy 85
chiral center 72 ion-exchange Protein Primary Structure and Evolution
enantiomers 72 chromatography 86 Andersson, L., Blomberg, L., Flegel, M., Lepsa, L., Nllsson, B.,
absolute [Link] 74 size-exclusion & Verlander, M. (2000) Large-scale synthesis of peptides. Biopoly-
chromatography 87 mers 55, 227-250.
D, L system 74
A discussion of approaches to manufacturing peptides as
polarity 74 affinity chromatog-
pharmaceuticals.
absorbance, A 76 raphy 88
Dell, A. & Morris, H.R. (2001) Glycoprotein structure determina-
zwitterion 78 high-performance liquid tion by mass spectrometry. Science 291, 2351-2356.
isoelectric pH (isoelectric chromatography Glycoproteins can be complex; mass spectrometry is a pre-
point, pi) 80 (HPLC) 88 ferred method for sorting things out.
peptide 82 electrophoresis 88 Delsuc, F., Brinkmann, H., & Philippe H. (2005) Phylogenomics
protein 82 sodium dodecyl sulfate and the reconstruction of the tree of life. Nat. Reu. Genet. 6, 361-375.
peptide bond 82 (SOS) 89 Gogarten, J.P. & Townsend, J.P. (2005) Horizontal gene transfer,
oligopeptide 82 isoelectric focusing 90 genome innovation and evolution. Nat. Reu. Microbiol. 3, 679--687.
polypeptide 82 primary structure 92 Gygi, S.P. & Aebersold, R. (2000) Mass spectrometry and pro-
secondary struc- teomics. Curr: Opin. Chem. Biol. 4, 489-494.
oligomeric protein 84
Uses of mass spectrometry to identify and study cellular proteins.
protomer 84 ture 92
tertiary structure 92 Koonin, E.V., Tatıısov, R.L., & Galperin, M.Y. (1998) Beyond
coııjugated protein 84
complete genomes: from sequence to structure and function. Curr.
prosthetic group 84 quaternary struc- Opin. Struct. Biol. 8, 355-363.
crude extract 85 ture 92 A good discussion about the possible uses of the increasing
fraction 85 Edman degradation 95 amount of information on protein sequences.
fractionation 85 proteases 95 Li, W.-H. & Graur, D. (2000) Fundamentals of Molecular Evotu-
dialysis 85 proteome 100 tion, 2nd edn, Sinauer Associates, lnc., Sunderland, MA.

[, 08] Amino Acids, Peptides, and Proteins

. tely titrated (second equivalence


. A ,-eıy readable text describing metlıods used to anıılyze pro- (i) Glycine ıs comp1e
teın aııd nuclcic acid seqııences. Chapter 6 provides one of Uıe best
point). . . + N CH -coo-
G) The predominant specıes ıs H3 ~ 2
a,-ailable descriptions of how evolutionary trees are constructed from .
sequence dala · ıs -1 .
(k) The average net charge ofgly cıne .
(1) Glycine is present predOmın
Mann, M. & Wllm, M. (1995) Electrospray mass spectrometry for . antly as a 50:50 mıxture of
_
protein characterization. 'lhmds Biochem. Sci. 20, 219-224.
+H3N-CH2-COOH and + ~3N-:-CH2-COO .
An approachable sıınınıary of this technique for beginners.
(m) This is the isoelectnc ~oın~-
Mayo, K.8. (2000) Recent advances in the design and construction
of synthetic peptides: for the love of basics or just for the technology (n) This is the end of the tıtration. .
of il 'lhmds BiotechnoL 18, 212-217. (o) These are the worst pH regions for buffenng power.
Miranda, L.P. & Alewood, P.F. (2000) Challenges for protein
chemical synthesis in the 21st century: bridging genomics [Link] pro-
teomics. Biopalymers 55, 217-226.
,......-------,-7
12 Jl-,3Q ___________________ ,(V)

This and Mayo, 2000 (above), describe how to make peptides


and splice them together to address a wide raııge of problems in pro- 10 - ~-,!iQ____ _
tein biochemistry.
Rokas, A., W'ılllams, B.L., Kiııg, N., & Carroll, S.B. (2003) 8
Geııome-scale approaches to resolving incongrueııce in molecular
phylogeııies. Nature 425, 798-804.
How sequence comparisons of mııltiple proteins can yield accu- pH 6 _&,~L-- --- (ill)
rate evolutionary information.
Sanger, F. (1988) Sequeııces, sequences, sequences. Annıı. Rev. 4
Biochem. 57, 1-28.
Anice historical account of the development of sequencing
methods. 2
Snel, B., Buynen, M.A., & Dutilh B.E. (2005) Genome trees and
the natııre of geııome evolutioıı.Annıı. Rev. [Link] 59, 191-209. o (I)
0.5 1.0 1.5 2.0
Zuckerkandl, E. & Paııling, L. (1965) Molecııles as docııments of
evolutionaıy history. J. Theor: BioL 8, 357-366. OW (equivalents)
Many consider this the foundiııg paper in the fıeld of molecıılar
evolution.
3. How Much Alanine Is Present as the Completely
Uncharged Species? Ata pH equal to the isoelectric point o:
Problems - - - - - - - - - alanine, the net charge on alanine is zero. Two structures can
be drawn that have a net charge of zero, but the predominant
1. Absolute Configuration of Citrıılline
The citrulline fonn of alanine at its pl İS zwitterionic.
isolated from watennelons has the structure shown below. Is it
a D· or L-amino acid? Explain. CHa o CHa o
+ ~ 1 1~
H3N-C-C H2N-C-cz_
CH2(CH 2)zNH-C-NH2
l
H-C--NH3
f
+ il
O
* "o-
Zwitterionic
il
Uncharged
OH

coo-
(a) Why is alanine predominantly zwitterionic rather thaıı
2. Relationship between the 1itration Curve and the
completely uncharged at its pi?
Acid-Base Properties of Glycine A 100 mL solution of
(b) What fraction of alanine is in the completely un-
0.1 M glycine at pH 1.72 was titrated with 2 M NaOH solution.
charged fonn at its pi? Justify your assumptions.
The pH was monitored and the results were plotted as shown
in the fo!lowirıg graph. The key points in the titration are des- 4. lonization [Link] ofHistidine Each ionizable group of an
ignated I to V. For each of the statements (a) to (o), identify amino acid can exist in one of two states, charged or neutral.
the appropriate key point in the titration and justify your The electric charge on the functional group is detennined by
choice. the relationship between its pKa and the pH of the solution.
(a) Glycine is present predominantly as the species This relationship İS described by the Henderson-Hasselbalch
+H;ıN--CHz-COOH. equation.
(b) The average net charge of glycine is +!. (a) Histidine has three ionizable functional groups. Write
(c) Half of the amino groups are ionized. the equilibrium equations for its three ionizations and assign
(d) The pH is equal to the pKa of the carboxyl group. the proper pKa for each ionization. Draw the structure of histi-
(e) The pH is equal to the pKa of the protonated amino dine in each ionization state. What is the net charge on the his-
group. tidine molecule in each ionization state?
(f) Glycine has its maximum buffering capacity. (b) Draw the structures of the predominant ionization
(g) The average net charge of glycine is zero. state ofhistidine at pH 1, 4, 8, and 12. Note that the ionization
(h) The carboxyl group has been completely titrated (first state can be approximated by treating each ionizable group
equivalence point) . independently.
Problems [109]

? (c) What is the_ne~ c~3:1"8e of histidine at pH 1, 4, 8, and


12. For each pH, will histıdine mıgra·te t 8. The Size of Proteins What is the approximate molecular
oward the anode ( +) weight of a protein with 682 amino acid residues in a sing]e
or cathode ( - ) when placed in an electric fıeld?
polypeptide chain?
5. Separation .of Aınino Acids by lon-Exchange Chro-
9. The Number of Tryptophan Residues in Bovine
mato~aphy Mix~ures of amino acids can be analyzed by fırst
Serum Albumin A quantitative amino acid analysis reveals
separatıng the mıxture into its components through ion-
exchange chromatography. that bovine serum albumin (BSA) contains 0.58% tryptophan
. Amino acids placed on a catıon-
.
(M, 204) by weight.
exchange re sın ( see Fig. 3-17a) conta· . ulf
- ınıng s onate (a) Calculate the minimum molecular weight of BSA (i.e.,
(-SOa ) groups flow down the column at different rates be-
assume there is only one Trp residue per protein molecule).
cause of two factors that influence their movement· (1) . .
. b . ıonıc (b) Gel filtration of BSA gives a molecular weight estimate
attractıon etween the sulfonate residues on th
.. e co1umn and of 70,000. How many Trp residues are present in a molecule of
posıtively charged functional groups on the amino acid d serum albumin?
(2) hydrophobic interactions between amino acid side ~~s
and the stro~ hydrophobic backbone of the polystyrene resin. 10. Subunit Cornposition ofa Protein A protein has a mo-
For each paır of amino acids listed, determine which will be lecular mass of 400 kDa when measured by gel filtration. When
eluted fırst from the cation-exchange column by a pH 7_0 buffer. subjected to gel electrophoresis in the presence of sodium dode-
(a) Asp and Lys cyl sulfate (SDS), the protein gives three bands with molecular
(b) Arg and Met masses of 180, 160, and 60 kDa. When electrophoresis is carried
(c) Glu and Val out in the presence of SDS and dithiothreitol, three bands are
(d) Gly and Leu again formed, this time with molecular masses of 160, 90, and 60
(e) Ser and Ala kDa. Determine the subunit composition of the protein.

6. Naming the Stereoisomers of lsoleucine The struc- 11. Net Electric Charge of Peptides A peptide has the
ture of the amino acid isoleucine is sequence

Glu-His-Trp-Ser-Gly- Leu- Arg-Pro-Gly


coo-
+ 1
H3N-C-H (a) What is the net charge of the molecule at pH 3, 8, and
1 11? (Use pKa values for side chains and terminal amino and
H- C- CH
1 3 carboxyl groups as given in Table 3-1.)
CH2 (b) Estimate the pl for this peptide.
6H 3 12. lsoelectric Point of Pepsin Pepsin is the name given to
(a) How many chiral centers does it have? a mix of several digestive enzymes secreted (as larger precur-
sor proteins) by g!ands that !ine the stomach. These g!ands
(b) How many optical isomers?
also secrete hydrochloric acid, which dissolves the particulate
(c) Draw perspective fonnulas for ali the optical isomers
matter in food, allowing pepsin to enzymatically cleave indi-
of isoleucine.
vidual protein molecules. The resulting mixture of food, HCI,
7. Comparing the pKa Values of Alanine and Polyala- and digestive enzymes is known as chyme and hasa pH near 1.5.
nine The titration curve of alanine shows the ionization of What pl would you predict for the pepsin proteins? What func-
two functional groups with pKa values of 2.34 and 9.69, corre- tional groups must be present to confer this pl on pepsin? Which
sponding to the ionization of the carboxyl and the protonated amino acids in the proteins would contribute such groups?
amino groups, respectively. The titration of di-, tri-, and larger
13. The lsoelectric Point of Histones Histones are pro-
oligopeptides of alanine also shows the ionization of only two
teins found in eukaryotic celi nuclei, tightly bound to DNA,
functional groups, although the experimental pKa values are
which has many phosphate groups. The pl of histones is very
different. The trend in pKa values is summarized in the table.
high, about 10.8. What amino acid residues must be present in
Amino acid or peptide pKl pK2 relatively large numbers in histones? In what way do these
residues contribute to the strong binding of histones to DNA?
Ala 2.34 9.69
14. Solubility of Polypeptides üne method for separating
Ala- Ala 3.12 8.30
polypeptides makes use of their different solubilities. The sol-
Ala-Ala- Ala 3.39 8.03 ubility of large polypeptides in water depends on the relative
Ala- (Ala)n-Ala, n 2:4 3.42 7.94 polarity of their R groups, particularly on the number of ion-
ized groups: the more ionized groups there are, the more solu-
(a) Draw the structure of Ala- Ala-Ala. Identify the func- ble the polypeptide. Which of each pair of the polypeptides
tional groups associated with pKı and pK2. that follow is more soluble at the indicated pH?
(b) Why does the value of pK1 increase with each addi- (a) (GIY)20 or (Glu)20 at pH 7.0
tional Ala residue in the oligopeptide? (b) (Lys-Ala) 3 or (Phe-Met) 3 at pH 7.0
(c) Why does the value of pK2 decrease with each addi- (c) (Ala-Ser- Gly) 5 or (Asn-Ser- His) 5 at pH 6.0
tional Ala residue in the oligopeptide? (d) (Ala-Asp-Gly) 6 or (Asn-Ser- His) 5 at pH 3.0
_____-_-_-_-_-_-_-_-_-_-_----- ----~~ ~~- -- ~---
o]
[ 11 Amino Acids, Pepıides, and Proıeins

mimic some of the properties of opiates. Some researchers


15. Puriflcation of an Enzyme A biochemist discovers and
consider these peptides to be the brain's own painkillers. Using
purifies a new enzyme, generating the purification lable below.
the infonnation below, determine the amino acid sequence of
the opioid leucine enkephalin. Explain how your structure is
Total
protein Activity consistent with each piece of infonnation.
Procedııre (mg) (onits) (a) Complete hydrolysis by 6 M HCI at 110 °C followed by
amino acid analysis indicated the presence of Giy, Leu, Phe,
l. Crude extract 20,000 4,000,000 and Tyr, in a 2:1:1:1 molar ratio.
2. Precipitation (salt) 5,000 3,000,000 (b) Treatment of the peptide with 1-fluoro-2,4-
3. Precipitation (pH) 4,000 1,000,000 dinitrobenzene followed by complete hydrolysis and chro-
4. Ion-exchange 200 800,000 matography indicated the presence of the 2,4-dinitrophenyJ
chromatography derivative of tyrosine. No free tyrosine could be found.
5. Affınity 50 750,000
(c) Complete digestion of the peptide with chymotrypsin
chromatography
675,000 followed by chromatography yielded free tyrosine and leucine,
6. Size-exclusion 45
chromatography plus a tripeptide containing Phe and Giy in a 1:2 ratio.
19. Structure of a Peptide Antibiotic from Bacillus
brevis Extracts from the bacterium Bacillus brevis contain
(a) From the infomıation given in the table, calculate the
a peptide with antibiotic properties. This peptide forms com-
specific activity of the enzyme after each purification procedure.
plexes with metal ions and seems to disrupt ion transport
(b) Which of the purification procedures used for this en-
across the celi membranes of other bacterial species, killing
zyme is most effective (i.e., gives the greatest relative increase
in purity)? them. The structure of the peptide has been determined from
(c) Which ofthe purification procedures is least effective? the following observations.
(a) Complete acid hydrolysis of the peptide followed by
(d) Is there any indication based on the results shown in
amino acid analysis yielded equimolar amounts of Leu, Orn,
the table that the enzyme after step 6 is now pure? What else
could be done to estimate the purity ofthe enzyme preparation? Phe, Pro, and Yal. Om is ornithine, an amino acid not pres-
ent in proteins but present in some peptides. it has the
16. Dialysis A purified protein is in a Hepes (N-(2-hydroxy- structure
ethy!)piperazine-N'-(2-ethanesulfonic acid)) buffer at pH 7 with
H
500 mM NaCL A sarnple (1 mL) of the protein solution is placed + 1
in a tube made of dialysis membrane and dialyzed against 1 L of H3N-CH2-CH 2-CH2-y-coo -
the sarne Hepes buffer with OmM Nacı. Small molecules and ions +NH3
(such as Na+, cı-, and Hepes) can diffuse across the dialysis
membrane, but the protein cannot. (b) The molecular weight of the peptide was estimated as
(a) ünce the dialysis has come to equilibrium, what is the about 1,200.
concentration of NaCI in the protein sample? Assume no vol- (c) The peptide failed to undergo hydrolysis when treated
ume changes occur in the sarnple during the dialysis. with the enzyme carboxypeptidase. This enzyme catalyzes the
(b) If the original 1 mL sample were dialyzed twice, succes- hydrolysis of the carboxyl-terminal residue of a polypeptidc
sively, against 100 mL of the sarne Hepes buffer with OmM NaCI, unless the residue is Pro or, for some reason, does not contain
what would be the final NaCI concentration in the sample? a free carboxyl group.
(d) Treatment of the intact peptide with l-fluoro-2,4-
17. Peptide Purification At pH 7.0, in what order would the dinitrobenzene, followed by complete hydrolysis and chro-
following three peptides be eluted from a colurnn filled with a matography, yieldı::d only free amino acids and the following
cation-exchange polymer? Their amino acid compositions are: derivative: ·
Protein A: Ala 10%, Glu 5%, Ser 5%, Leu 10%, Argl0%,
His 5%, ile 10%, Phe 5%, Tyr 5%, Lys 10%, Giy 10%, Pro 5%,
and Trp 10%.
Protein B: Ala 5%, Yal 5%, Giy 10%, Asp 5%, Leu 5%, Arg
5%, ile 5%, Phe 5%, Tyr 5%, Lys 5%, Trp 5%, Ser 5%, Thr 5%,
Glu 5%, Asn 5%, Pro 10%, Met 5%, and Cys 5%.
Protein C: Ala 10%, Glu 10%, Giy 5%, Leu 5%, Asp 10%,
(Hint: The 2,4-dinitrophenyl derivative involves the amino
Arg 5%, Met 5%, Cys 5%, Tyr 5%, Phe 5%, His 5%, Yal 5%, Pro
group ofa side chain rather than the a-amino group.)
5%, Thr 5%, Ser 5%, Asn 5%, and Gln 5%.
(e) Partial hydrolysis of the peptide followed by chro-
18. Sequence Determination of the Brain Peptide mat_ogra~hic separation and sequence analysis yielded the fol-
Leucine Enkephalin A group of peptides that influence lowıng di- and tripeptides (the amino-tenninal amino acid is
nerve transmission in certain parts of the brain has been iso- always at the left):
lated from normal brain tissue. These peptides are known as
Leu-Phe Phe-Pro Om-Leu Val-Om
opioids, because they bind to specific receptors that also bind
opiate drugs, such as morphine and naloxone. Opioids thus Val- Orn- Leu Phe-Pro-Val Pro-Val-Orn
Problems [111]
Given the above infonnation, deduce the amino .
of the peptide antibiotic. Show your reasonı· Whacıd sequence buffered to a pH of 7.2. Why do you use beef heart tissue
· d t
amve a a structure demonstrate th t I·t .
ng. en you have
. and ~n such large quantity? What is the purpose of keepirı:ı
. ' a ıs consıstent with
each experımental observation. the tıssue cold and suspending it in 0.2 M sucrose, at pH
7.2? What happens to the tissue when it is homogeni,zed?
20. Efflciency in Peptide Sequencın'g A . . (b) You subject the resulting heart homogenate, which is
. peptıde wıth the
pnmaryd sbtructure Lys-Arg-Pro-Leu-Ile-Asp-Gly-Ala is se- dense and opaque, to a series of differential centrifugation
quence Ythe Edman procedure· If each Edman cycle ıs . steps. What does this accomplish?
. 96%
efficıent, what percentage of the amın·0 a 'd lib . (c) You proceed with the purification using the super-
. cı s erated ın the
fourth cycle will be leucine?· Do the calculat·ıon a second time natant fraction that contains mostly intact mitochondria. Next
but assume a 99% efficiency for each cycle. ' you osmotically lyse the mitochondria. The lysate, which is
less dense than the homogenate, but stili opaque, consists pri-
21. Sequence Comparisons Proteın· s cailed molecular
marily of mitochondrial membranes and internal mitochondr-
chaperones
. .(described in Chapter 4) assı·st ın
. the process of
ial contents. To this lysate you add ammonium sulfate, a high]y
proteın folding. üne class of chaperone found ın .
. orgarusms
. soluble salt, to a specific concentration. You centrifuge the so-
from bactena to mammals is heat shock protein 90 (Hsp90). lution, decant the supematant, and discard the pellet. To the
All Hsp90
,, chaperones
. contain a 10 amino acı'd "sıgna
. t ure se-
supernatant, which is clearer than the lysate, you add more
. . which allows for ready identification of these pro-
quence, ammonium sulfate. ünce again, you centrifuge the sample, but
teıns ın sequence databases. Tvvo representations of this this time you save the pellet because it contains the citrate
signature sequence are shown below.
synthase. What is the rationale f or the two-step addition of
the salt? ·

~f~EiErRS'
(d) You solubilize the ammonium sulfate pellet containing
the mitochondrial proteins and dialyze it overnight against
large volumes ofbuffered (pH 7.2) solution. Why isn't ammo-
nium sulfate included in the dialysis bujfer? Why do you
1 2 3 4 5 6 7 8 9 10 use the bujfer solution instead of water?
N C (e) You run the dialyzed solution over a size-exclusion
(a) in this sequence, which amino acid residues are invari- chromatographic column. Following the protocol, you collect
ant (conserved across all species)? thefirst protein fraction that exits the column and discard the
(b) At which position(s) are amino acids limited to those fractions that elute from the column later. You detect the pro-
with positively charged side chains? For each position, which tein by measuring UV absorbance (at 280 nm) by the fractions.
amino acid is more commonly found? What does the instruction to collect the first fraction tell
you about the protein? Why is UV absorbance at 280 nm a
(c) At which positions are substitutions restricted to
good way to monitor f or the presence of protein in the
amino acids with negatively charged side chains? For each po-
elutedfractions?
sition, which amino acid predominates?
(f) You place the fraction collected in (e) on a cation-
(d) There is one position that can be any amino acid, al-
exchange chromatographic column. After discarding the initial
though one amino acid appears much more often than any other.
solution that exits the column (the flowthrough), you add a
What position is this, and which amino acid appears most often?
washing solution of higher pH to the column and collect the
22. Biochemistry Protocols: Your First Protein Purifl- protein fraction that immediately elutes. Explain what you
cation As the newest and least experienced student in a bio- are doing.
chemistry research lab, your first few weeks are spent washing (g) You run a small sample of your fraction, now very re-
glassware and labeling test tubes. You then graduate to making duced in volume and quite clear (though tinged pink), on an
buffers and stock solutions for use in various laboratory proce- isoelectric focusing gel. When stained, the gel shows three
dures. Finally, you are given responsibility far purifying a pro- sharp bands. According to the protocol, the citrate synthase is
tein. It is citrate synthase (an enzyme of the citric acid cycle, the protein with a pl of 5.6, but you decide to do one more as-
to be discussed in Chapter 16), which is located in the mito- say of the protein 's purity. You cut out the pl 5. 6 band and sub-
chondrial matrix. Following a protocol for the purifıcation, you ject it to SOS polyacrylamide gel electrophoresis. The protein
proceed through the steps below. As you work, a more experi- resolves as a single band. Why were you unconpinced of the
enced student questions you about the rationale for each purity of the "single" protein band on your isoelectric fo-
procedure. Supply the answers. (Hint: See Chapter 2 for infor- cusing gel? What did the results of the SDS gel tell you?
mation about osmolarity; see p. 7 for information on separation Why is it important to do the SDS gel electrophoresis after
of organelles from cells.) the isoelectricfocusing?
(a) You pick up 20 kg of beef hearts from a nearby slaugh-
terhouse (muscle cells are rich in mitochondria, which supply Data Analysis Problem - - -- - - -
energy for muscle contraction). You transport the hearts on
ice, and perform each step of the purifıcation on ice or in a 23. Determining the Amino Acid Sequence of Insulin
walk-in cold room. You homogenize the beef heart tissue in a Figure 3- 24 shows the amino acid sequence of the honnone in-
high-speed blender in a medium containing 0.2 M sucrose, sulin. This structure was determined by Frederick Sanger and
[ 11 ~ Amino Acids, Peptides, and Proteins

his coworkers. Most of lhis work is described in a series of arti- 6. Isolated four of Uıe DNP-peptides, which were named 81
clcs published in Uıe Biochemical Journal from 1945 to 1955. through B4.
\Vhen Sanger and colleagues began their work in 1945, it 7. Strongly hydro]yzed each DNP-peptide to give free amino
was kno"ıı that insulin was a small protein consisting of two or acids.
four polypeptide chains linked by disulfide bonds. Sanger and 8. ldentified the amino acids in each peptide with paper
his coworkers had developed a few simple metlıods for study-
chromatography.
ing protein sequences.
Treatrnent with FDNB. FDNB (l-fluoro-2,4-dinitroben- The results were as follows:
zene) reacted with free amino (but not amido or guanidino) Bl: a-DNP-phenylalanine only
groups in proteins to produce dinitrophenyl (DNP) derivatives B2: a-DNP-phenylalanine; valine
of amino acids: B3: aspartic acid; a-DNP-phenylalanine; valine
B4: aspartic acid; glutamic acid; a-DNP-phenylalanine;
valine
(c) Based on these <lata, what are the fırst four (amino-
terminal) amino acids of the B chain? Explain your reasoning.
Amine FDNB DNP-amine (d) Does this result match the known sequence of insulin
(Fig. 3-24)? Explain any discrepancies.
Acili Hydrolysi.s. Boiling a protein with 10% HCl for sev- Sanger and colleagues used these and related methods to
eral hours hydro]yzed ali of its peptide and amide bonds. Short determine the entire sequence of the A and B chains. Their se-
treatments produced short po]ypeptides; the longer the treat- quence for the A chain was as follows (amino terminus on left):
ment, the more complete Uıe breakdown of the protein into its
amino acids. 1 5 10
Oxidation of Cysteines. Treatment ofa protein with per- Gly-Ile-Val-Glx-G!x-Cys-Cys-Ala-Ser-Val-
15 20
formic acid cleaved ali the disulfide bonds and converted ali
Cys residues to cysteic acid residues (Fig. 3-26). Cys-Ser-Leu-Tyr-Glx-Leu-Gl.x-Asx-Tyr-Cys-Asx
Paper Chromatography. This more primitive version of
Because acid hydrolysis had converted ali Asn to Asp and ali
thin-layer chromatography (see Fig. 10-24) separated com-
Gln to Glu, these residues had to be designated Asx and Glx,
pounds based on their chemical properties, allowing identifi-
respectively (exact identity in the peptide unknown). Sanger
cation of single amino acids and, in some cases, dipeptides.
Thin-layer chromatography also separates Iarger peptides. solved this problem by using protease enzymes that cleave
As reported in his first paper (I 945), Sanger reacted in- peptide bonds, but not the amide bonds in Asn and Gln
sulin with FDNB and hydrolyzed the resulting protein. He residues, to prepare short peptides. He then determined the
found many free amino acids, but only three DNP-amino number of amide groups present in each peptide by measuring
acids: a-DNP-glycine (DNP group attached to the a-amino the NH! released when the peptide was acid-hydrolyzed.
group); a-DNP-phenylalanine; and e-DNP-lysine (DNP at- Some of the results for the A chain are shown below. The pep-
tached to the e-amino group). Sanger interpreted these results tides may not have been completely pure, so the numbers
as showing that insulin had two protein chains: one with Giy at were approximate-but good enough for Sanger's purposes.
its amino terminus and one with Phe at its amino terİninus.
Peptide Peptide Noınber of amide
üne of the two chains also contained a Lys residue, not at the
name sequence groups in peptide
amino terminus. He named the chain beginning with a Giy
residue "A" and the chain beginning with Phe "B." Acl Cys-Asx 0.7
(a) Explain how Sanger's results support his conclusions. Apl5 Tyr-Glx-Leu 0.98
(b) Are the results consistent with the known structure Apl4 Tyr-Glx-Leu-Glx 1.06
ofinsulin (Fig. 3-24)? Ap3 Asx-Tyr-Cys-Asx 2.10
In a later paper (1949), Sanger described how he used Apl Glx-Asx-Tyr-Cys-Asx 1.94
these techniques to determine the first few amino acids Ap5pal Gly-Ile-Val-Glx 0.15
(amino-terminal end) of each insulin chain. To analyze the B Ap5 Gly-Ile-Val-Glx-Glx-Cys- Cys-
chain, for example, he carried out the following steps: Ala-Ser-Val- Cys-Ser-Leu 1.16

1. Oxidized insulin to separate the A and B chains.


2. Prepared a sample of pure B chain with paper (e) Based on these data, determine the amino acid se-
chromatography. quence of the A chain. Explain how you reached your answer.
Compare it with Figure 3-24.
3. Reacted the B chain with FDNB.
References
4. Gently acid-hydrolyzed the protein so that some small
peptides would be produced. Sanger, F. (1945) The free amino groups of insıılin. Biochem. J 39
507-515. . '
5. Separated the DNP-peptides from the peptides that did Sanger, F. (1949) The terminal peptides of insıılin. Bi<xhem. J 45
not contain DNP groups. 563-574. '

You might also like