International Journal of Advanced Engineering & Applications, ISSN: 0975 -7791 2010
Prediction of Protein Secondary Structure based on GOR Algorithm
integrating with multiple sequences alignment
1
Binod Kumar, Assistant Professor, Computer Sc. Dept, ISTAR, Vallabh Vidyanagar , Anand, Gujarat, India.
2
Dr. N. N. Jani, Director, Faculty of Computer Sc. & Information Tech., Kadi Sarva Vishvavidyalaya Univ.,
Gandhinagar, India
1
[Link].1970@[Link],2drnnjanicsd@[Link]
ABSTRACT GOR (Garnier, Osguthorpe, and Robson) [2]
method is based on the assumption that amino
We have discussed in this paper a new method for acids neighboring the central amino acid
the prediction of the protein secondary structure residue influence the secondary structure. The
from the amino acid sequence. The method is GOR method uses principles of information
based on the most recent version of the standard theory to derive predictions. In GOR method,
GOR algorithm. A significant improvement is three scoring matrices, containing in each
obtained by combining multiple sequence
column the probability of finding each amino
alignments with the GOR method. Additional
improvement in the predictions is obtained by a acid at one of the 17 positions, are prepared.
simple correction of the results when helices or One matrix corresponds to the central (eighth)
sheets are too short, or if helices and sheets are amino acid being found in a α helix, the
direct neighbors along the sequence .The second for the amino acid being in a β strand,
imposition of the requirement that the prediction the third a coil, and the fourth, a turn. A
must be strong enough, i.e. that the difference candidate sequence is analyzed by each of the
between the probability of the predicted (most three to four matrices by a sliding window of
probable) state and the probability of the second 17 residues. Each matrix is positioned along a
most probable state must be larger than a certain candidate sequence and the matrix giving the
minimum value also improves significantly highest score predicts the structural state of
secondary structure predictions
the central amino acid. At least 4 residues in a
KEYWORDS: GOR algorithm; protein secondary row have to be predicted as an α helix and 2 in
structure; secondary structure prediction; multiple a row for a β strand for a prediction to be
sequence alignment. validated.
INTRODUCTION Matrix values are calculated in the same
manner as amino acid substitution matrices, in
Since structural information can provide that matrix values are calculated as log odds
insight into protein function, high-accuracy units representing units of information. The
prediction of protein structure from its information available as to the joint
sequence is highly desirable. The availability occurrence of secondary structural
of structural information may advance drug conformation S and amino acid a is given by
design efforts and provide a more detailed [3]:
understanding of protein-protein interaction
networks. Secondary structure prediction I (S; a) = log [P(S | a) / P(S)] [1]
methods are also useful for motif detection in
membrane proteins or for enhancing where P(S | a) is the conditional probability of
homology modeling [1]. conformation S given residue a, and P(S) is
the probability of conformation S. By Bayes’
102
International Journal of Advanced Engineering & Applications, ISSN: 0975 -7791 2010
rule, the probability of conformation S given + log [1 - P(S)/ P(S)] [6]
amino acid a, (S | a) is given by:
from which the following ratio of the joint
P(S | a) = P(S, a) / P(a) [2] probability of conformation Sm given a1,..aX to
where P(S, a) is the joint probability of S and the joint probability of any other conformation
a and P (a) is the probability of a. These may be calculated
probabilities can be estimated from the
frequency of each amino acid found in each P(Sm,a1,..aX)/[1–P(Sm,a1,..aX)]
structure and the frequency of each amino acid = {P(S)/ [1–P(S)]} e-I(ΔSm; a1,..aX) [7]
in the structural database. Given these
frequencies, Searching for all possible patterns in the
structural database would require an enormous
I (S; a) = log (fS,a / fS ) [3] number of proteins. Hence, three simplifying
approaches have been taken. First, it was
where fS,a is the frequency of amino acid a in assumed in earlier versions of GOR that there
conformation S and fS is the frequency of all is no correlation between amino acids in any
amino acid residues found to be in of the 17 positions (both the neighboring 8
conformation S. positions and the central amino acid position),
or that each amino acid position had a separate
The GOR method maximizes the information and independent influence on the structural
available in the values of fS,a and avoids data conformation of the central amino acid. The
size and sampling variations by calculating the steps are then: (1) Values for I (ΔS; a) in
information difference between the competing Equation 7 are calculated for each of the 17
hypotheses that residue a is in structure S, I positions; (2) these values are summed to
(S;a), or that a is in a different conformation approximate the value of I (ΔSm; a1,..aX) in
(not S), I (not S;a). This difference I (ΔS;a) is Equation 6; (3) the probability ratios in
calculated from Equation 3 with simple Equation 7 are calculated.
substitutions by
APPLICATIONS AND EXPERIMENT
I (ΔS;a) = I (S; a) - I (not S; a)
= log {P(S,a)/[1- P(S,a)]}+log{[1- P(S)/ P(S)]}[4] Here some experiments on amino acid
properties ofprotein sample have been
which is derived from the observed amino performed. Finally (Figure 3) probabilities of
acid data as: secondary structure elements by GOR has
been performed taking consideration of
I (ΔS; a) = log [fS,a / (1 - fS,a)] + log [ (1 - fS )/ fS[5] Sequence index, Amino acid type, Helix
probability, Sheet probability, Coil probability
where the frequency of finding amino acid a and GOR prediction.
not in conformation S is 1 - fS,a and of not
finding any amino acid in conformation S is Table 1: Calculation of Statistics of protein properties
1 - fS. Equation 4 is used to calculate the
information difference for a series of x
consecutive positions neighboring sequence
position m,
I (ΔSm; a1,..aX)=log[P(Sm,a1,..aX)/(1 -P(Sm,a1,..aX)]
103
International Journal of Advanced Engineering & Applications, ISSN: 0975 -7791 2010
Table 1: Calculation of Statistics of protein properties
(continue)
Figure 1: Pot of Histogram of amino acid properties of
proteins in parallel
Figure 2: Predicts protein secondary structure using
GOR methods [4]
104
International Journal of Advanced Engineering & Applications, ISSN: 0975 -7791 2010
Figure 3: Predicts coiled coil regions in protein
sequence [5]
Figure3: Probabilities of secondary structure elements by GOR
For sequence no. 1 to 30 For sequence no. 31 to 60
1.2 1.2
1 1
0.8 0.8
Coil probability Coil probability
0.6 Sheet probability 0.6 Sheet probability
Helix probability Helix probability
0.4 0.4
0.2 0.2
0 0
CCCCCCCCEE EECCCCCE E EECCCCCCCCC CCCCCHHHHHHHHHHEEEECCCCCCCEEEE
MNG T EGP N F Y V P F S N A T G V VR S P F E Y PQ Y Y L A E PWQ F S M L A A Y M F L L I V L G F P I N F L T L Y
1 2 3 4 5 6 7 8 9 101112131415161718192021222324252627282930 313233343536373839404142434445464748495051525354555657585960
For sequence no. 61 to 90 For sequence no. 91 to 120
105
International Journal of Advanced Engineering & Applications, ISSN: 0975 -7791 2010
1.2 1.2
1 1
0.8 0.8
Coil probability Coil probability
0.6 Sheet probability 0.6 Sheet probability
Helix probability Helix probability
0.4 0.4
0.2 0.2
0 0
EHHHHHHCCCCCCHHHHHHHHHHHHHHHCC CCEEEEECCCCEEEECCCCCCCCCEEEECC
V T VQH K K L R T P L N Y I L L N L A V A D L F MV L GG F T ST L YT SL HGYF VFGPTGCNL EGF F AT LG
616263646566676869707172737475767778798081828384858687888990 919293949596979899100101102103104105106107108109110111112113114115116117118119120
For sequence no. 121 to 150 For sequence no. 151 to 180
1.2 1.2
1 1
0.8 0.8
Coil probability Coil probability
0.6 Sheet probability 0.6 Sheet probability
Helix probability Helix probability
0.4 0.4
0.2 0.2
0 0
CCHHHHHHHHHHHHHHHHHHCCCCCCCCCC CEEEEEEE EEEEEEECCCCCCCCCCCCCCC
G E I A LWS L V V L A I E R Y V V V C K PMS N F R F G E N H A I MG V A F T W V M A L A C A A P P L A GW S R Y I P
121122123124125126127128129130131132133134135136137138139140141142143144145146147148149150 151152153154155156157158159160161162163164165166167168169170171172173174175176177178179180
neighboring region, or of a neighboring amino
acid and the central one, influence the
RESULT AND DISCUSSION conformation of the central one. Thus, there
are 17 x 16/2 = 136 possible pairs to use for
It has been observed [6] that in new versions frequency measurements and to examine for
of GOR is that certain pair-wise combinations correlation with the conformation of the
of an amino acid in the neighboring region central residue [7]. With the advent of a large
and central amino acid influence the number of protein structures, it has become
conformation of the central amino acid. This possible to assess the frequencies of amino
requires a determination of the frequency of acid combinations and to use this information
amino acid pairs between each of the 16 for secondary structural predictions. For α
neighboring positions and the central one, helices [8], a prediction is made when four of
both for when the central residue is in six amino acids have a high probability >1.03
conformation S and when the central residue is of being in an α helix. For β strands, the
not in conformation S. presence in a sequence of three of five amino
acids with a probability of >1.00 of being in a
Finally, in the most recent version of GOR, β strand predicts a nucleation event for a β
the assumption is made that certain pair-wise strand. These nucleated regions are extended
combinations of amino acids in the along the sequence in each direction until the
106
International Journal of Advanced Engineering & Applications, ISSN: 0975 -7791 2010
prediction values for four amino acids drops 8. Rychlewski L, Fischer D. LiveBench:
below 1. If both α-helical and β-strand regions continuous benchmarking of prediction
are predicted, the higher probability prediction servers. [Link]://[Link]/LiveBench.
is used.
CONCLUSION
We have shown that the GOR prediction
algorithm based on information theory and
incorporating multiple sequence alignment
information is quite successful in its accuracy
of secondary structure prediction. The GOR
method predicts 64% of the residue
conformations in known structures and quite
drastically (36.5%) under predicts the number
of residues in β strands.
The GOR method benefits from its relative
simplicity and low computational resource
requirements, which makes it possible to do
predictions in real time without long waits for
results. The GOR method predicts the
probabilities of the three conformational states
for each residue in the sequence, and this
information can be used for further analyses or
simulations.
REFERENCES
1. [Link]
2. Davit W. Mount .Bioinformatics,
Sequence and Genome Analysis. Gold
Spring Harbor Laboratory Press; 427-440.
3. [Link]
4. [Link]
5. [Link]
/[Link]
6. [Link]://ScienceDirect-Polymer Protein
secondary structure prediction based on
the GOR algorithm incorporating multiple
sequence alignment [Link]
7. Frishman D, Argos P. Seventy-five percent
accuracy in protein secondary structure
prediction. Proteins 1997; 27:329–335
107