Biochemistry Unit3 Part2
Biochemistry Unit3 Part2
Disulfide bond _J
1 (cystine) ..J
7
TH2SH NH O=C
1 1
CHOH HC-CH2-S-S-CH2-CH
1 1 . 1
CHOH C=O HN
1
CH2SH f 7ı
Dithlothreitol (D'IT)
\_ r1 f
NH O O O=C ~NH O=C
1 11 1 1
HC-CH2-s-o- -o-i-cH2-6H HC-CH2-SH HS-CH2-CH
l 11 il 1 1 1
C=O O O HN C=O HN~
,J Cysteic acid
7ı (
ı
residues carboxyınethylation
by
TABLE 3-7
Reagent (biological source)* Cleavage pointi
Trypsin (bovine pancreas) Lys, Arg (C)
SubmaxiUarus protease (mouse submaxillary gland) Arg (C)
Chymotrypsin (bovine pancreas) Phe, Trp, Tyr (C)
St,a,phylococcus aureus V8 protease (bacterium S. aureus) Asp, Glu (C)
Asp-N-protease (bacterium Pseudomonas fragi) Asp, Glu (N)
Pepsin (porcine stomach) Leu, Phe, Trp, Tyr (N)
Endoproteinase Lys C (bacteriuı:n Lysobacter erızymogenes) Lys (C)
Cyanogen bromide Met (C)
•Ali reagents except cyanogen bromide are proteases. Ali are available from commerclal sources.
tResidues furnishing the primary recognition point for the protease or reagent; peptide bond cleavage occurs on either the car-
bonyt (C) or the amino (N) side of the indicated amino acid residues,
Arnong proteases, the ctigestive enzyme trypsin cat- terminal Lys or Arg. The fragments produced by trypsin
alyzes the hydrolysis of only those peptide bonds in (or other enzyme or chernical) action are then sepa-
which the carbonyl group is contributed by either a Lys rated by chromatographic or electrophoretic methods.
oran Arg residue, regardless of the length or amino acid
sequence of the chain. The nurnber of smaller peptides Seqnencing the Peptides Each peptide fragrnent re-
produced by trypsin cleavage can thus be precticted sulting from the action of trypsin is sequenced separately
frorn the total nurnber of Lys or Arg residues in the by the Edman procedure.
original polypeptide, as determined by hydrolysis of an
intact sample (Fig. 3-27). A polypeptide with three Ordering the Peptide Fragments The order of the
Lys and/or Arg residues (as in Fig. 3-27) wi1l usually "trypsin fragments" in the original polypeptide chain
yield four srnaller peptides on cleavage with trypsin. must now be deterrnined. Another sample of the intact
Moreover, all except one of these wi1l have a carboxyl- polypeptide is cleaved into fragments using a dif[erent
3.4 The Structure of Proteins: Primary Structure [ 97]
enzyme or reagent, one that cleaves peptide bonds at amino acid has been identified before the original cleavage
points other th8:Il those cleaved by trypsin. For exarnple, of the protein, this information can be used to establish
cyanogen bromıde cleaves only those peptide bonds in which fragment is derived from the amino terminus. The
which the carbonyl group is contributed by Met. The two sets of fragments can be compared for possible errors
fragments resulting from this second procedure are in deterrnining the amino acid sequence of each fragment.
then separated and sequenced as before. If the second cleavage procedure fails to establish continu-
The amino acid sequences of each fragment obtained ity between all peptides from the first cleavage, a third or
by the two cleavage procedures are examined, with the even a fourth cleavage method must be used to obtain a set
objective of finding peptides from the second procedure of peptides that can provide the necessaıy overlap(s).
whose sequences establish continuity, because of overlaps,
between the fragments obtained by the first cleavage pro- Locating Disulfide Bonds If the primary structure
cedure (Fig. 3-27). Overlapping peptides obtained from the includes disulfıde bonds, their locations are determined
second fragmentation yield the correct order of the peptide in an additional step after sequencing is completed. A
fragments produced in the first. If the amino-terminal sarnple of the protein is again cleaved with a reagent
8 YLIACGPMTK
terminus because it
does not end with
R (Arg) or K (Lys).
eatablish
sequence
Amino
s Carboxyl
EGAAYHDFEPIDPRGASMALIKYLIACGPMTKDCVHSD terminus
terminus
FIGURE 3-27 Cleaving proteins and sequencing and ordering the thus only one possibility for location of the disulfide bond. in polypep-
peptide fragments. First, the amino acid composition and amino- tides with three or more Cys residues, the position of disulfide bonds
terminal residue of an intact sample are determined. Then any disulfide can be determined as described in the text. (The one-letter symbols for
bonds are broken before fragmenting so that sequencing can proceed amino acids are given in Table 3-1.)
efficiently. in this example, there are only two Cys (C) residues and
[98] Amlno Acids, P,ptid,s, and Proteins
·:LL
M = n 2 [(mlzh - X]
This calculation using the mlz values for any two peaks 50+
in a spectrum such as that shown in Figure ı b usually l 100 !
provides the mass of the protein (in this case, aerolysin ~
k; 47,342 Da) with an error of only ±0.01%. Generating -~ 75 ,... 40+ 47,000 48,000
.s ! Mr
several sets of peaks, repeating the calculation, and av- .s
CI)
eraging the results generally provides an even more ac- :oaı> 50 ,... 30+
curate value for M. Computer algorithms can transforrn !
the mlz spectrum into a single peak that also provides a ~
25 ,...
very accurate mass measurement (Fig. lb, inset).
Mass spectrometry can also be used to sequence
short stretches of polypeptide, an application that has o
.,,til.
1
w
1 1 1
J..1 1
J ıt
emerged as an invaluable tool for quickly identifying un- 800 1,000 1,200 1,400 1,600
lrnown proteins. Sequence inforrnation is extracted using (b) mlz
a technique called tandem MS, or MS/MS. A solution FIGURE 1 Electrospray mass spectrometry of a protein. (a) A protein
containing the protein under investigation is first treated solution is dispersed into highly charged droplets by passage through a
with a protease or chemical reagent to hydrolyze it to a needle under the influence ofa high-voltage electric field. The droplets
mixture of shorter peptides. The mixture is then injected evaporate, and the ions (with added protons in this case) enter the
into a device that is essentially two mass spectrometers mass spectrometer for m/z measurement. The spectrum generated (b)
in tandem (Fig. 2a, top). In the first, the peptide mixture is a family of peaks, with each successive peak (from right to left) cor-
is sorted and the ionized fragments are manipulated so responding to a charged species increased by 1 in both mass and
that only one of the several types of peptides produced charge. lnset: a computer-generated transformation of this spectrum.
by cleavage emerges at the other end. The sample of the
selected peptide, each molecule of which has a charge
somewhere along its length, then travels through a vac- The second mass spectrometer then measures the
uum chamber between the two mass spectrometers. In mlz ratios of ali the charged fragments (uncharged
this collision celi, the peptide is further fragmented by fragments are not detected). This generates one or
high-energy impact with a "collision gas," a small amount more sets of peaks. A given set of peaks (Fig. 2b) con-
of a noble gas such as helium or argon that is bled into sists of ali the charged fragments that were generated
the vacuum chamber. This procedure is designed to frag- by breaking the same type of bond (but at different
ment many of the peptide rnolecules in the sample, with points in the peptide) and are derived frorn the same
each individual peptide broken in only one place, on av- side of the bond breakage, either the carboxyl- or
erage. Most breaks occur at peptide bonds. This frag- amino-terminal side. Each successive peak in a given set
mentation does not involve the addition of water (it is has one less amino acid than the peak before. The dif-
done in a near-vacuum), so the products rnay include ference in mass from peak to peak identifies the amino
rnolecular ion radicals such as carbonyl radicals (Fig. 2a, acid that was lost in each case, thus revealing the se-
bottorn). The charge on the original peptide is retained quence of the peptide. The only ambiguities involve
on one of the fragments generated from it. leucine and isoleucine, which have the same mass.
(conti nued on next page)
[1 oo] Amino Acids, Peptides, and Proteins
- METHODS
MS-1
Collision
celi MS-2 Detector
FIGURE 2 Obtaining protein sequence information with tandem MS.
(a) After proteolytic hydrolysis, a protein solution is injected into a mass
spectrometer (MS-1 ). The different peptides are sorted so that only one
~ ~ @) type is selected for further analysis. The selected peptide is further frag-
Electrospray Separation Breakage
;, ru,atioo, 1 mented in a chamber between the two mass spectrometers, and m/z for
each fragment is measured in the second mass spectrometer (MS-2).
Many of the ions generated during this second fragmentation result
from breakage of the peptide bond, as shown. These are called b-type
or y-type ions, depending on whether the charge is retained on the
amino- or carboxyl-terminal side, respectively. (b) A typical spectrum
R1 O R8 O R6 with peaks representing the peptide fragments generated from a sample
H H I b
1 il il ~ H H 1 ~O
11,N-C-C-N -C-C-N- C-C N-C-C-N- C-C of one small peptide (1 Oresidues). The labeled peaks are y-type ions.
H H I il H H I il H 'o-
R2 O ~ O The large peak next to y5" is a doubly charged ion and is not part of the
y
y set. The successive peaks differ by the mass ofa particular amino acid
in the original peptide. in this case, the deduced sequence was
R1 O R8 O RI
1 il H H I il H H 1 ~O Phe-Pro-Gly-Gln--(lle/Leu)-Asn-Ala-Asp-(lle/Leu)-Arg. Note the am-
RıN-C-C-N-C-C-N-C-C• •N-C-C-N-C -C
H H I il H H I il H 'o- biguity about ile and Leu residues, because they have the same molec-
Rı O R4 O ular mass. in this example, the set of peaks derived from y-type ions
(a) predominates, and the spectrum is greatly simplified as a result. This is
because an Arg residue occurs at the carboxyl terminus of the peptide,
and most of the positive charges are retained on this residue.
100 . ..
ı
Y2" fragments can be unambiguously distinguished from that
~
l
75 ·I
consisting of the arnino-ternıinal fragments. Because th~
b
·; Ya" l bond breaks generated between the spectrometers (in the
.ı
ı::
.s 50 collision cell) do not yield full [Link] and amino groups
.s y/ Ys" .j at the sites of the breaks, the only intact a-amino arıd
CI) 11·
.::... 25
Ys°
j a-carboxyl groups on the peptide fragments are those at
<il
~ 1 the very ends (Fig. 2a). The two sets of fragments ca.11
o thereby be identified by the resulting slight differences in
200 400 600
(b) mlz mass. The amino acid sequence derived from one set can
be confirmed by the other, improving the confidence in the
sequence information obtained.
Even a short sequence is often enough to pennit un-
The charge on the peptide can be retained on either ambiguous association of a protein with its gene, if the
the [Link]- or arnino-ternıinal fragment, and bonds other gene sequence is known. Sequencing by mass spectrom-
tlıan the peptide bond can be broken in the [Link] etry cannot replace the Edman degradation procedure
process, with the result that multiple sets of peaks are usu- for the sequencing of long polypeptides, but it is ideal
alJy generated. The two most prominent sets generally for proteornics research aimed at cataloging the hun-
consist of charged fragments derived from breakage of the dreds of cellular proteins that rnight be separated on a
peptide bonds. The set consisting of the [Link]-ternıinal two-dimensional gel.
Amino acid
Fmoc residue
-t
Insoluble
Fmoc ·- R1 il
-N- CH- O
_ Attachment of carboxyl-terminal
by Fmoc group C- 0 (D amino acid to reactive
,J H
I group on resin.
...1 , ·-- · , cı-
R1 O
ı Fmoc ~ N-6H-~-O-CH2 l\.____r\ ., ___________ _____\
1
_. ,..,- H ______... ~ - 1
R2 O ,._.ı. . r·
ı
1 il
Fmoc - N- CH- c-o- Protecting group is removed
1 ® by fluahing with solution
H containing a mild organic base.
, " R 2
O
Q NH ©
4
a-Aminogroupofamino
acid 1 attacks activated
carboxyl group of amino acid
1
1
1
1
1
1
1
. 1 1
6
11 1 2 to form peptide bond.
Fmoc -N- CH- C-0- C 1
1
1 il
o~-~-~-o
1
o 1
1
- H 1
1
1
1
H H 1
Dicyclohexylurea byproduct 1
1
1
1
~ Q ~ Ü •~~® to@ 1
l-0-0'
Completed peptide is
HF deprotected 88 in
@ reaction®; HF cleaves
ester linkage between
peptide and resin.
R2 O R' O ı ,
+ CH-
H N- 1 11
C-N- 1
CH- cil - o - + F-CH 2
R. Bruce Merrifield 3 1 -
1921-2 006 H
[1 oı] Amino Acids, P,ptides, and Proteins
-
3.4 The Structure of Proteins: Primary Structure
~ :b
Ol
=-Cli
234567s
four posıtıons, followed by an invariant G and an invari-
ant K. The last position is either S or T.
S~quence logos provide a more infonnative and
(a) N C grap~c representation of an amino acid (or nucleic acid)
D-{W}-IDNSJ-{ILVFYW}-IDENSTGJ-IDNQGHRKJ-{ GP}- multıple sequence alignment. Each logo consists of a
4--------~=~-~
3
[LIVMCJ-IDENQSTAGCJ-x(2)-IDEI- ILIVMFYwJ. stack of symbols for each position in the sequence. The
overall height of the stack (in bits) indicates the degree
of sequence conservation at that position while the
-~ 2 height of each symbol in the stack indicates the relative
ıı:ı
1 fr:quency of that amino acid (or nucleotide). For amino
o 1 acıd sequences, the colors denote the characteristics of
2 3 4 5 6 7 8 9 10 11 12 13 the amino acid: polar (G, S, T, Y, C, Q, N) green; basic (K,
(b) N
C R, H) blue; acidic (D, E) red; and hydrophobic (A V L I
FIGURE 1 Representations of two consensus sequences. (a) p loop, an P, W, F, M) black. The classifıcation of amino acid~ ~ t~
ATP-binding structure; (b) EF hand, a Caı+-binding structure. scheme is somewhat different from that in Table 3- ı and
Figure 3-5. The amino acids with aromatic side chains
In one type of consensus sequence designation
are subsumed into the nonpolar (F, W) and polar (Y)
(shown at the top of (a) and (b)), each position is sepa-
classifıcations. Glycine, always hard to group, is assigned
rated from its neighbor by a hyphen. A position where
to the polar group. Note that when multiple amino acids
any amino acid is allowed is designated x. Ambiguities are acceptable at a particular position, they rarely occur
are indicated by listing the acceptable amino acids for a with equal probability. üne or a few usually predominate.
given position between square brackets. For example, in The logo representation makes the predominance clear,
(a} [AG] means Ala or Gly. If all but a few amino acids and a corıserved sequence in a protein is made obvious.
are allowed at one position, the amino acids that are not However, the logo obscures some amino acid residues
allowed are listed between curly brackets. For example, that may be allowed at a position, such as the Cys that
in (b) (W} means any amino acid except Trp. Repetition occasionally occurs at position 8 of the EF hand in (b).
and macromolecular structures, has given rise to the The field of molecular evolution is often traced to
new field of bioinformatics. üne outcome of this disci- Emile Zuckerkandl and Linus Pauling, whose work in
pline is a growing suite of computer programs, many the mid-1960s advanced the use of nucleotide and pro-
readily available on the Internet, that can be used by any tein sequences to explore evolution. The premise is de-
scientist, student, or knowledgeable layperson. Each ceptively straightforward. If two organisms are closely
protein's function relies on its three-dimensional struc- related, the sequences of their genes and proteins
ture, which in turn is determined largely by its primary should be similar. The sequences increasingly diverge as
structure. Thus, the biochemical information conveyed the evolutionary distance between two organisms in-
by a protein sequence is limited only by our own under- creases. The promise of this approach began to be real-
standing of structural and functional principles. The ized in the 1970s, when Carı Woese used ribosomal RNA
constantly evolving tools of bioinformatics make it sequences to define the Archaea as a group of living
possible to [Link] functional segments in new_proteins organisms distinct from the Bacteria and Eukarya (see
and help establish both their sequence and their struc- Fig. 1- 4). Protein sequences offer an opportunity to
tural relationships to proteins already in the databases. greatly refine the available information. With the advent
On a different level of inquiry, protein sequences are of genome projects investigating organisms from bacte-
beginning to tell us how the proteins evolved and, ulti- ria to humans, the number of available sequences is
mately, how life evolved on this planet. growing at an enormous rate. This information can be
◄
[ 104] Amino Acids, Peptides, and Proteins
used. to trace biological history. The challenge is in are called orthologs. The process of tracing evolution
leanung to read the genetic hieroglyphics. involves fırst identifying suitable families of homolo-
~~olution has not taken a sirnple linear path. Corn- gous proteins and then using thern to reconstruct evo-
plexıties abound in any atternpt to mine the evolution- lutionary paths.
aıy ~orrnation stored in protein sequences. Fora given Hornologs are identifıed through the use of increas-
proteın, the amino acid residues essential for the activ- ingly powerful computer prograrns that can directıy
ity of the protein are conserved over evolutionary time. cornpare two or rnore chosen protein sequences, or can
The residues that are less important to function rnay search vast databases to .find the evolutionary relatives
vary over tirne-that is, one arnino acid rnay substitute of one selected protein sequence. The electronic search
for another-and these variable residues can provide process can be thought of as sliding one sequence past
the inforrnation to trace evolution. Arnino acid substitu- the other until a section with a good rnatch is found.
tions are not always randorn, however. At sorne posi- Within this sequence aligrunent, a positive score is as-
tions in the prirnaıy structure, the need to rnaintain signed for each position where the arnino acid residues
protein function rnay rnean that only particular arnino in the two sequences are identical-the value of the
acid substitutions can be tolerated. Some proteins have score varying frorn one program to the next-to provide
more variable amino acid residues than others. For a measure of the quality of the alignrnent. The process
these and other reasons, different proteins can evolve at has some cornplications. Sornetimes the proteins being
düferent rates. compared rnatch well at, say, two sequence segments,
Another complicating factor in tracing evolutionary and these segments are connected by less related se-
history is the rare transfer of a gene or group of genes quences of düferent lengths. Thus the two rnatching
from one organism to another, a process called lateral segments cannot be aligned at the sarne time. To handle
gene transfer. The transferred genes may be quite this, the computer program introduces "gaps" in one of
similar to the genes they were derived frorn in the origi- the sequences to bring the rnatching segments into
nal organism, whereas rnost other genes in the same two register (Fig. 3-30). Of course, ifa suffıcient nurnber of
organisms may be quite distantly related. An example of gaps are introduced, alrnost any two sequences could be
lateral gene transfer is the recent rapid spread of brought into sorne sort of aligrunent. To avoid uninfor-
antibiotic-resistance genes in bacterial populations. The rnative alignments, the prograrns include penalties for
proteins derived from these transferred genes would not each gap introduced, thus lowering the overall align-
be good candidates for the study of bacterial evolution, rnent score. With electronic trial and error, the program
because they share only a very limited evolutionary his- selects the alignment with the optimal score that rnaxi-
tory with their "host" organisms. rnizes identical amino acid residues while minirnizing
The study of molecular evolution generally focuses the introduction of gaps.
on families of closely related proteins. In most cases, Identical arnino acids are often inadequate to identify
the families chosen for analysis have essential func- related proteins or, rnore importantly, to deterrnine how
tions in cellular metabolism that must have been pres- closely related the proteins are on an evolutionary tµne
ent in the earliest viable cells, thus greatly reducing scale. Arnore useful analysis includes a consideration of
the chance that they were introduced relatively re- the chernical properties of substituted arnino acids. When
cently by lateral gene transfer. For example, a protein arnino acid substitutions are found within a protein fam-
called EF-la (elongation factor la) is involved in the ily, rnany of the düferences rnay be conservative-that
synthesis of proteins in ali eukaryotes. A similar pro- is, an arnino acid residue is replaced by a residue havmg
tein, EF-Tu, with the sarne function, is found in bacte- similar chernical properties. For exarnple, a Glu residue
ria. Similarities in sequence and function indicate that may substitute in one farnily rnernber for the Asp residue
EF-la and EF-Tu are members ofa family of proteins found in another; both arnino acids are negatively
that share a common ancestor. The members of protein charged. Such a conservative substitution should logi-
families are called homologous proteins, or ho- cally garner a higher score in a sequence alignrnent than
mologs. The concept of a homolog can be further re- does a nonconservative substitution, such as the replace-
fined. If two proteins in a family (that is, two ment of the Asp residue with a hydrophobic Phe residue.
homologs) are present in the same species, they are re- For most efforts to find homologies and explore evo-
ferred to as paralogs. Homologs from düferent species lutionary relationships, protein sequences (derived either
FIGURE 3-30 Aligning protein sequences with the use of gaps. Shown bacterial species, E. co/i and Bacillus subtilis. lntroduction ofa gap in ıhe
here is the sequence alignment of a short section of the Hsp70 proteins B. subtilis sequence allows a better alignment of amino acid residues on
(a widespread class of protein-folding chaperones) from two well-studied either side of the gap. ldentical amino acid residues are shaded.
3.4 The Structure of Proteins: Primary Structure [1 os]
Signature sequence
Archaea Halobacterium halobium IGHVD HGK S TMVGR LYET GSVPEHVIEQH
Sulfolobus solfataricus IGHVaHUK S TLVGR L LMDRGFIDEKT KEA
Eukaryotes { Saccharomyces cerevisiae IGHVDSGK S~ TTGH L IYKC GGIDKRTIEKF
.. . Homo sapiens IGHVD S G~S TTTGH L IYKC GGIDKRTIEKF
Gram-posıtive bactenum Bacillus subtilis IGHVD HGK STMVGR ITTV
Gram-negative bacterium Escherichia coli IGHVD HGKTTLTAA ITTV
FIGURE3-31 A signature sequence in the EF-1a/EF-Tu protein family. although the sequences of the insertions are quite distinct for the two
The signature sequence (boxed) is a 12-residue insertion near ıhe groups. The variation in the signature sequence reflects the significant
amino terminus of the sequence. Residues that align in all species are evolutionary divergence that has occurred at this site since it first ap-
shaded yellow. Both archaea and eukaryotes have the signature, peared in a common ancestor of both groups.
directly from protein sequencing or from the sequencing ture sequences for the group in which they are found.
of the DNA encoding the protein) are superior to non- An example of a signature sequence is an insertion of
genic nucleic acid sequences (those that do not encode a 12 amino acids near the amino terminus of the EF-
protein or functional RNA). For a nucleic acid with its la/EF-Tu proteins in all archaea and eukaryotes but not
four different types of residues, random alignment' of non- in bacteria (Fig. 3-31). This particular signature is one
homologous sequences will generally yield matches for at of many biochemical clues that can help establish the
least 25% of the positions. Introduction of a few gaps can evolutionary relatedness of eukaryotes and archaea.
often increase the fraction of matched residues to 40% or Other signature sequences allow the establishment of
more, and the probability of chance alignment of unre- evolutionary relationships among groups of organisms
lated sequences becomes quite high. The 20 different at many different taxonomic levels.
amino acid residues in proteins greatly lower the probabil- By considering the entire sequence of a protein,
ity of uninformative chance alignrnents of this type. researchers can now construct more elaborate evolu-
The programs used to generate a sequence align- tionary trees with many species in each taxonomic
ment are complemented by rnethods that test the relia- group. Figure 3-32 presents one such tree for bacte-
bility ofthe alignments. A cornmon computerized test is ria, based on sequence divergence in the protein
to shuf:fle the amino acid sequence of one of the pro- GroEL (a protein present in ali bacteria that assists in
teins being compared to produce a random sequence, the proper folding of proteins). The tree can be re-
then to instruct the program to align the shuf:fled se- fined by basing it on the sequences of rnultiple pro-
quence with the other, unshuf:fled one. Scores are as- teins and by supplementing the sequence information
signed to the new alignment, and the shuffling and with data on the unique biochemical and physiologi-
aligrunent process is repeated many times. The original cal properties of each species. There are many meth-
aligrunent, before shnffling, should have a score signifi- ods for generating trees, each method with its own
cantly higher than any of those within the distribution advantages and shortcomings, and many ways to rep-
of scores generated by the random alignments; this in- resent the resulting evolutionary relationships. In
creases the confidence that the sequence alignment has Figure 3-32, the free end points of lines are called
identifıed a pair of homologs. Note that the absence of "external nodes"; each represents an extant species,
a significant alignment score does not necessarily mean and each is so labeled. The points where two lines
that no evolutionary relationship exists between two come together, the "internal nodes," represent ex-
prnteins. As we shall see in Chapter 4, three-dimensional tinct ancestor species. in most representations (in-
structural similarities sornetirnes reveal evolutionary cluding Fig. 3-32), the lengths of the lines connecting
relationships where sequence hornology has been the nodes are proportional to the number of amino
wiped away by time. acid substitutions separating one species frorn an-
Use ofa protein family to explore evolution requires other. If we trace two extant species to a comrnon in-
the [Link] of family rnernbers with similar rnolec- ternal node (representing the common ancestor of
ular functions in the widest possible range of organisrns. the two species), the length of the branch connecting
Information frorn the family can then be used to trace each external node to the internal node represents
the evolution of those organisrns. By analyzing the se- the nurnber of amino acid substitutions separating
quence divergence in selected protein families, investi- one extant species from this ancestor. The sum of the
gators can segregate organisms into classes based on lengths of all the line segments that connect an extant
their evolutionary relationships. This inforrnation must species to another extant species through a cornmon
be reconciled with more classical examinations of the ancestor reflects the nurnber of substitutions separat-
physiology and biochemistry of the organisrns. ing the two extant species. To determine how much
Certain segments of a protein sequence may be time was needed for the various species to diverge,
found in the organisrns of one taxonomic group but not the tree must be calibrated by comparing it with in-
in other groups; these segments can be used as signa- formation from the fossil record and other sources.
[106] Amino Acids, Peptides, and Particles
Bacteroides [
[Link]
• [ Chlamydia trachomatis
Chlamydia psittaci
Porp/ıyroınoııa& giııgivalis
Borrelia bu'lldorferi l Spirochaetes
LtptMpira interrogaM
Oj
·c
r
[
Ltgioııellıı pneumophila
Ytnıinia enttrocolitica
Tl,~pJ,llk ...... ,,.,
Baci/luı subtilis
Staphylococcus aureus
Clostridium acetobutylicum
l
Jow
G+C
Oj
·c
...
Q)
"
Oj
.Q
~
Salmonella typhi CIMtridium per{ringen• Q)
Eschtrichia coli
~
Oj
.Q
·~---ı
0 1/J
~
Riı:kdtsia
/3 0
ı:l.
l tıutsugamu,hi
Mycobacterium leprae .
high
G+C ..~
Bradyrhizobiumjaponiı:um Mycobacterium tuberculosı& c:,
a Streptomyce• albu, [genel
Agrobacterium tumefaciena
Zymamanas mabilis
Cyanobacteria and
chloroplasts
FIGURE 3-32 Evolutionary tree derived from amino acid sequence vergence observed in the GroEL family of proteins. Also included in this
1
1
comparisons. A bacterial evolutionary tree, based on the sequence di- tree (lower right) are the chloroplasts (ehi.) of some nonbacterial species.
11
,,1 As more sequence information is made available in organism on Earth. The story is a work in progress, of
databases, we can generate evolutionary trees based course (Fig. 3-33). The questions being asked andan-
1 on multiple proteins. And we can refine these trees as swered are fundamental to how humans view themselves
1 additional genomic information emerges from increas- and the world around them. The fıeld of molecular evolu-
'
ingly sophisticated methods of analysis. All of this work tion promises to be among the most vibrant of the scien-
moves us toward the goal of creating a detailed tree of tifıc frontiers in the twenty-fırst century.
life that describes the evolution and relationship of every
LowG+C, Crenarchaeota
gram-positive Thermo- Desulfurococcales
proteales
Thermotogales Sulfolobales
Aquificales ~ Euryarchaeota
Spirochaetes / Halobacteriales
Chlamydiales~ \ ~ Methanosareinales
'• \ ,, Th I t1
De~ococcales .::: - - ~ ' ermop asma a e.s
High G + C,_ ~ - Bacter·a Arch'' Arehaeoglobales
gram-negatıve ı aea
Cyanobacteria ~ -~ , , Methanococcales
. Mıtochon drıa Thermococcales
Proteobactena '-.. Eukarya
. - - - -- Chloroplast~
[Land plantıJ \ ~-,.~- Opisthokonta
Green algae ~ - -
Planta~ı!:oı:~s
Mycetozoans
Pelobionts
1.1/Ğ
',~
'\ ',,~
/ \ ',, ----....::::---Fungi
Ch~:~:=~ates
(multicellular aniınals)
Radiolaria
Entamoebae Cereozoa
Amoebozoa Rhizari
Diplomonads Jalı. b"d Alveolates a
O1
Euglenoids c~to- Stramenopiles
Excavata phytes Haptophytes
Chromalveolata
FIGURE 3-33 Aconsensus tree of life. The tree shown here is based on The tree presents only a fraction of the available information, as well as
analyses of many different protein sequences and additional genomic only a fraction of the issues remaining to be resolved. Each exıanı
features. Branches shown as dashed lines remain under investigation. group shciwn is a complex evolutionary story unıo itself.
Further Reading [101]
SUMMARY 3.4 The Structure of Proteins: consensus sequence 102 homolog 104
bioinformatics 103 paralog 104
Primary Structure lateral gene transfer 104 ortholog 104
■ Differences in protein function result from homologous proteins 104 signature sequence 105
differences in amino acid composition and sequence.
Some variations in sequence are possible for a
particular protein, with little or no effect on function. Further Reading - - - - - - - - -
■ Amino acid sequences are deduced by fragmenting Amino Acids
po]ypeptides into smaller peptides with reagents Dougherty, D.A. (2000) Unnatural arnino acids as probes of protein
known to cleave specific peptide bonds; determining structure and function. Gurr: Opin. Ghem. Biol. 4, 645-652.
the amine acid sequence of each fragment by the Greenstein, J.P. & Wınitz, M. (1961) Ghemistry ojtheAmino
automated Edman degradation procedure; then Acids, 3 Vols, John Wıley & Sons, New York.
ordering the peptide fragments by finding sequence Krell, G. (1997) o-Arnino acids in animal [Link]. Rev.
overlaps between fragments generated by different Biochem. 66, 337-345.
reagents. A protein sequence can also be deduced Details the occurrence of these unusual stereoisomers of
arnino acids.
from the nucleotide sequence of its corresponding
gene in DNA. · Meister, A. (1965) Biochemistry of the Amino Acids, 2nd edn,
Vols 1 and 2, Academic Press, ine., New York.
■ Short proteins and peptides (up to about 100 Encyclopedic treatment of the properties, occurrence, and
residues) can be chemically synthesized. The metabolism of arnino acids.
peptide is built up, one amino acid residue at a time, Peptides and Proteins
while tethered to a solid support.
Creighton, T.E. (1992) Proteins: Structures and Molecular
■ Protein sequences are a rich source of information Properties, 2nd edn, W. H. Freeman and Company, New York.
about protein structure and function, as well as the Very useful general source.
evolution of life on Earth. Sophisticated methods Working with Proteins
are being developed to trace evolution by analyzing Dunn, M.J. & Corbett, J.M. (1996) 'I\vo-dimensional polyacryl-
the resultant slow changes in amine acid sequences amide gel electrophoresis. Methods Enzymol. 271, 177-203.
of hornologous proteins. A detailed description of the technology.
Komberg, A. (1990) Why purify enzyrnes? Methods Enzymol.
182, 1-5.
KeyTerms - - - - - - - - - - - The critical role of classical biochemical methods in a new age.
Turms in bold are defined in the glossary. Scopes, R.K. (1994) Protein Purijication: Principles and
Practice, 3rd edn, Springer-Verlag, New York.
. A good source for more complete descriptions of the principles
amino acids 72 column chromatog-
underlying chromatography and other rnethods.
Rgroup 72 raphy 85
chiral center 72 ion-exchange Protein Primary Structure and Evolution
enantiomers 72 chromatography 86 Andersson, L., Blomberg, L., Flegel, M., Lepsa, L., Nllsson, B.,
absolute [Link] 74 size-exclusion & Verlander, M. (2000) Large-scale synthesis of peptides. Biopoly-
chromatography 87 mers 55, 227-250.
D, L system 74
A discussion of approaches to manufacturing peptides as
polarity 74 affinity chromatog-
pharmaceuticals.
absorbance, A 76 raphy 88
Dell, A. & Morris, H.R. (2001) Glycoprotein structure determina-
zwitterion 78 high-performance liquid tion by mass spectrometry. Science 291, 2351-2356.
isoelectric pH (isoelectric chromatography Glycoproteins can be complex; mass spectrometry is a pre-
point, pi) 80 (HPLC) 88 ferred method for sorting things out.
peptide 82 electrophoresis 88 Delsuc, F., Brinkmann, H., & Philippe H. (2005) Phylogenomics
protein 82 sodium dodecyl sulfate and the reconstruction of the tree of life. Nat. Reu. Genet. 6, 361-375.
peptide bond 82 (SOS) 89 Gogarten, J.P. & Townsend, J.P. (2005) Horizontal gene transfer,
oligopeptide 82 isoelectric focusing 90 genome innovation and evolution. Nat. Reu. Microbiol. 3, 679--687.
polypeptide 82 primary structure 92 Gygi, S.P. & Aebersold, R. (2000) Mass spectrometry and pro-
secondary struc- teomics. Curr: Opin. Chem. Biol. 4, 489-494.
oligomeric protein 84
Uses of mass spectrometry to identify and study cellular proteins.
protomer 84 ture 92
tertiary structure 92 Koonin, E.V., Tatıısov, R.L., & Galperin, M.Y. (1998) Beyond
coııjugated protein 84
complete genomes: from sequence to structure and function. Curr.
prosthetic group 84 quaternary struc- Opin. Struct. Biol. 8, 355-363.
crude extract 85 ture 92 A good discussion about the possible uses of the increasing
fraction 85 Edman degradation 95 amount of information on protein sequences.
fractionation 85 proteases 95 Li, W.-H. & Graur, D. (2000) Fundamentals of Molecular Evotu-
dialysis 85 proteome 100 tion, 2nd edn, Sinauer Associates, lnc., Sunderland, MA.
◄
[, 08] Amino Acids, Peptides, and Proteins
coo-
(a) Why is alanine predominantly zwitterionic rather thaıı
2. Relationship between the 1itration Curve and the
completely uncharged at its pi?
Acid-Base Properties of Glycine A 100 mL solution of
(b) What fraction of alanine is in the completely un-
0.1 M glycine at pH 1.72 was titrated with 2 M NaOH solution.
charged fonn at its pi? Justify your assumptions.
The pH was monitored and the results were plotted as shown
in the fo!lowirıg graph. The key points in the titration are des- 4. lonization [Link] ofHistidine Each ionizable group of an
ignated I to V. For each of the statements (a) to (o), identify amino acid can exist in one of two states, charged or neutral.
the appropriate key point in the titration and justify your The electric charge on the functional group is detennined by
choice. the relationship between its pKa and the pH of the solution.
(a) Glycine is present predominantly as the species This relationship İS described by the Henderson-Hasselbalch
+H;ıN--CHz-COOH. equation.
(b) The average net charge of glycine is +!. (a) Histidine has three ionizable functional groups. Write
(c) Half of the amino groups are ionized. the equilibrium equations for its three ionizations and assign
(d) The pH is equal to the pKa of the carboxyl group. the proper pKa for each ionization. Draw the structure of histi-
(e) The pH is equal to the pKa of the protonated amino dine in each ionization state. What is the net charge on the his-
group. tidine molecule in each ionization state?
(f) Glycine has its maximum buffering capacity. (b) Draw the structures of the predominant ionization
(g) The average net charge of glycine is zero. state ofhistidine at pH 1, 4, 8, and 12. Note that the ionization
(h) The carboxyl group has been completely titrated (first state can be approximated by treating each ionizable group
equivalence point) . independently.
Problems [109]
6. Naming the Stereoisomers of lsoleucine The struc- 11. Net Electric Charge of Peptides A peptide has the
ture of the amino acid isoleucine is sequence
~f~EiErRS'
(d) You solubilize the ammonium sulfate pellet containing
the mitochondrial proteins and dialyze it overnight against
large volumes ofbuffered (pH 7.2) solution. Why isn't ammo-
nium sulfate included in the dialysis bujfer? Why do you
1 2 3 4 5 6 7 8 9 10 use the bujfer solution instead of water?
N C (e) You run the dialyzed solution over a size-exclusion
(a) in this sequence, which amino acid residues are invari- chromatographic column. Following the protocol, you collect
ant (conserved across all species)? thefirst protein fraction that exits the column and discard the
(b) At which position(s) are amino acids limited to those fractions that elute from the column later. You detect the pro-
with positively charged side chains? For each position, which tein by measuring UV absorbance (at 280 nm) by the fractions.
amino acid is more commonly found? What does the instruction to collect the first fraction tell
you about the protein? Why is UV absorbance at 280 nm a
(c) At which positions are substitutions restricted to
good way to monitor f or the presence of protein in the
amino acids with negatively charged side chains? For each po-
elutedfractions?
sition, which amino acid predominates?
(f) You place the fraction collected in (e) on a cation-
(d) There is one position that can be any amino acid, al-
exchange chromatographic column. After discarding the initial
though one amino acid appears much more often than any other.
solution that exits the column (the flowthrough), you add a
What position is this, and which amino acid appears most often?
washing solution of higher pH to the column and collect the
22. Biochemistry Protocols: Your First Protein Purifl- protein fraction that immediately elutes. Explain what you
cation As the newest and least experienced student in a bio- are doing.
chemistry research lab, your first few weeks are spent washing (g) You run a small sample of your fraction, now very re-
glassware and labeling test tubes. You then graduate to making duced in volume and quite clear (though tinged pink), on an
buffers and stock solutions for use in various laboratory proce- isoelectric focusing gel. When stained, the gel shows three
dures. Finally, you are given responsibility far purifying a pro- sharp bands. According to the protocol, the citrate synthase is
tein. It is citrate synthase (an enzyme of the citric acid cycle, the protein with a pl of 5.6, but you decide to do one more as-
to be discussed in Chapter 16), which is located in the mito- say of the protein 's purity. You cut out the pl 5. 6 band and sub-
chondrial matrix. Following a protocol for the purifıcation, you ject it to SOS polyacrylamide gel electrophoresis. The protein
proceed through the steps below. As you work, a more experi- resolves as a single band. Why were you unconpinced of the
enced student questions you about the rationale for each purity of the "single" protein band on your isoelectric fo-
procedure. Supply the answers. (Hint: See Chapter 2 for infor- cusing gel? What did the results of the SDS gel tell you?
mation about osmolarity; see p. 7 for information on separation Why is it important to do the SDS gel electrophoresis after
of organelles from cells.) the isoelectricfocusing?
(a) You pick up 20 kg of beef hearts from a nearby slaugh-
terhouse (muscle cells are rich in mitochondria, which supply Data Analysis Problem - - -- - - -
energy for muscle contraction). You transport the hearts on
ice, and perform each step of the purifıcation on ice or in a 23. Determining the Amino Acid Sequence of Insulin
walk-in cold room. You homogenize the beef heart tissue in a Figure 3- 24 shows the amino acid sequence of the honnone in-
high-speed blender in a medium containing 0.2 M sucrose, sulin. This structure was determined by Frederick Sanger and
[ 11 ~ Amino Acids, Peptides, and Proteins
his coworkers. Most of lhis work is described in a series of arti- 6. Isolated four of Uıe DNP-peptides, which were named 81
clcs published in Uıe Biochemical Journal from 1945 to 1955. through B4.
\Vhen Sanger and colleagues began their work in 1945, it 7. Strongly hydro]yzed each DNP-peptide to give free amino
was kno"ıı that insulin was a small protein consisting of two or acids.
four polypeptide chains linked by disulfide bonds. Sanger and 8. ldentified the amino acids in each peptide with paper
his coworkers had developed a few simple metlıods for study-
chromatography.
ing protein sequences.
Treatrnent with FDNB. FDNB (l-fluoro-2,4-dinitroben- The results were as follows:
zene) reacted with free amino (but not amido or guanidino) Bl: a-DNP-phenylalanine only
groups in proteins to produce dinitrophenyl (DNP) derivatives B2: a-DNP-phenylalanine; valine
of amino acids: B3: aspartic acid; a-DNP-phenylalanine; valine
B4: aspartic acid; glutamic acid; a-DNP-phenylalanine;
valine
(c) Based on these <lata, what are the fırst four (amino-
terminal) amino acids of the B chain? Explain your reasoning.
Amine FDNB DNP-amine (d) Does this result match the known sequence of insulin
(Fig. 3-24)? Explain any discrepancies.
Acili Hydrolysi.s. Boiling a protein with 10% HCl for sev- Sanger and colleagues used these and related methods to
eral hours hydro]yzed ali of its peptide and amide bonds. Short determine the entire sequence of the A and B chains. Their se-
treatments produced short po]ypeptides; the longer the treat- quence for the A chain was as follows (amino terminus on left):
ment, the more complete Uıe breakdown of the protein into its
amino acids. 1 5 10
Oxidation of Cysteines. Treatment ofa protein with per- Gly-Ile-Val-Glx-G!x-Cys-Cys-Ala-Ser-Val-
15 20
formic acid cleaved ali the disulfide bonds and converted ali
Cys residues to cysteic acid residues (Fig. 3-26). Cys-Ser-Leu-Tyr-Glx-Leu-Gl.x-Asx-Tyr-Cys-Asx
Paper Chromatography. This more primitive version of
Because acid hydrolysis had converted ali Asn to Asp and ali
thin-layer chromatography (see Fig. 10-24) separated com-
Gln to Glu, these residues had to be designated Asx and Glx,
pounds based on their chemical properties, allowing identifi-
respectively (exact identity in the peptide unknown). Sanger
cation of single amino acids and, in some cases, dipeptides.
Thin-layer chromatography also separates Iarger peptides. solved this problem by using protease enzymes that cleave
As reported in his first paper (I 945), Sanger reacted in- peptide bonds, but not the amide bonds in Asn and Gln
sulin with FDNB and hydrolyzed the resulting protein. He residues, to prepare short peptides. He then determined the
found many free amino acids, but only three DNP-amino number of amide groups present in each peptide by measuring
acids: a-DNP-glycine (DNP group attached to the a-amino the NH! released when the peptide was acid-hydrolyzed.
group); a-DNP-phenylalanine; and e-DNP-lysine (DNP at- Some of the results for the A chain are shown below. The pep-
tached to the e-amino group). Sanger interpreted these results tides may not have been completely pure, so the numbers
as showing that insulin had two protein chains: one with Giy at were approximate-but good enough for Sanger's purposes.
its amino terminus and one with Phe at its amino terİninus.
Peptide Peptide Noınber of amide
üne of the two chains also contained a Lys residue, not at the
name sequence groups in peptide
amino terminus. He named the chain beginning with a Giy
residue "A" and the chain beginning with Phe "B." Acl Cys-Asx 0.7
(a) Explain how Sanger's results support his conclusions. Apl5 Tyr-Glx-Leu 0.98
(b) Are the results consistent with the known structure Apl4 Tyr-Glx-Leu-Glx 1.06
ofinsulin (Fig. 3-24)? Ap3 Asx-Tyr-Cys-Asx 2.10
In a later paper (1949), Sanger described how he used Apl Glx-Asx-Tyr-Cys-Asx 1.94
these techniques to determine the first few amino acids Ap5pal Gly-Ile-Val-Glx 0.15
(amino-terminal end) of each insulin chain. To analyze the B Ap5 Gly-Ile-Val-Glx-Glx-Cys- Cys-
chain, for example, he carried out the following steps: Ala-Ser-Val- Cys-Ser-Leu 1.16