Protein Folding
and
Protein Threading
Protein threading
Structure is better conserved than sequence
Structure can adopt a
wide range of mutations.
Physical forces favor
certain structures.
Number of folds is limited.
Currently ~700
Total: 1,000 ~10,000 TIM barrel
Tolga Can, METU, CENG 465
Protein Threading
• Basic premise
The number of unique structural (domain) folds in nature
is fairly small (possibly a few thousand)
• Statistics from Protein Data Bank (~35,000 structures)
90% of new structures submitted to PDB in the past
three years have similar structural folds in PDB
Tolga Can, METU, CENG 465
Concept of Threading
o Thread (align or place) a query protein sequence
onto a template structure in “optimal” way
o Good alignment gives approximate backbone
structure
Query sequence
MTYKLILNGKTKGETTTEAVDAATAEKVFQYANDNGVDGEWTYTE
Template set
Tolga Can, METU, CENG 465
Protein Threading – energy function
MTYKLILNGKTKGETTTEAVDAATAEKVFQYANDNGVDGEWTYTE
how preferable to put
two particular residues
nearby: E_p how well a residue fits
a structural
environment: E_s
alignment gap
penalty: E_g
total energy: E_p + E_s + E_g
find a sequence-structure alignment
to minimize the energy function
Tolga Can, METU, CENG 465
Prediction of Protein Structures
• Examples – a few good examples
actual predicted actual predicted
actual predicted actual predicted
Tolga Can, METU, CENG 465
Prediction of Protein Structures
• Not so good example
Tolga Can, METU, CENG 465
CASP/CAFASP
• CASP: Critical CASP
Assessment of Predictor
Structure Prediction
• CAFASP: Critical
Assessment of Fully CAFASP
Automated Structure Predictor
Prediction
1. Won’t get tired
2. High-throughput
Tolga Can, METU, CENG 465
Protein Threading
Kristen Huber, UMass, EC 697S
Protein Threading
Kristen Huber, UMass, EC 697S
Protein Threading
❑ Jinbo Xu, Ying Xu, Dongsup Kim, Ming Li.
RAPTOR: Optimal Protein Threading by Linear Programming
Journal of Bioinformatics and Computational Biology, April 2003
❑ Given a query sequence S = (s1, s2, s3, …sn) and a template (library)
sequence T = (t1, t2, t3, …tm), pair up elements from S and T, by possibly
inserting gaps, while minimizing an energy function
❑ Assumptions
➢ the template is a sequence of cores (conserved segments – α-helix or β-
sheet) connected by loops
➢ gaps are allowed only within the loops
➢ only interactions between residues in the cores
are considered;
➢ interaction between residues is assumed to exist
if they are within 7 Ǻ and at least 4 positions away
Protein Threading
❑ Steps of the RAPTOR algorithm:
➢ build a contact map for the template structure
➢ find all possible alignments for each core within the query sequence
➢ build a contact map for the query sequence and template structure
➢ define energy function and carry out minimization
E =WmEm +Ws Es +W p E p +Wg Eg +Wss Ess
Em – mutation score
Es – environment fitness score
Ep – pairwise interaction score
Eg – gap penalty score
Ess – secondary structure compatibility
Wx – weights (determined experimentally)
Applications to Protein Contact maps
12 A Ca contact map for 2csn, casein
kinase-1
• contact maps specify
both secondary
structure and inter-
residue contacts
• a detailed contact map
provides sufficient
information to
reconstruct a 3-D
structure
• generation of a large
set of feasible contact
maps can reproduce
near native structures
(Smith et al, 1997)
Protein structure prediction: Contact maps
Some issues with distance based contact maps:
• typically use Ca-Ca distances
- dependent on appropriate choice of cutoff
• short cutoff distance biases map towards contacts within
secondary structures
• longer cutoff distance results in more contacts, and a noisy
data set
• Ca atoms in close proximity may have little interaction
- e.g. n, n+2 residues in an alpha helix
• contact with solvent not readily integrated
A tessellation procedure based on residue sidechains can
circumvent some of these issues
Voronoi Contact maps
• similar to Ca-Ca distance-based maps
• uses a tesselation procedure to determine if residues are in
contact
• contacts can be subdivided by type:
- sidechain contacts
- backbone contacts
- both sidechain and backbone
• results in recognizable patterns of interaction for within and
between secondary structures
• it is possible to integrate solvent contact into this scheme
as well
Voronoi Contact maps
Voronoi Contact maps
Voronoi map Ca distance map
Voronoi Contact map feature recognition
Using the contact preferences from residue-residue scores, it is possible
to recognize regions of secondary structure, and interactions between
secondary structure elements:
alpha helix:
alpha-alpha:
antiparallel
beta-beta
beta-alpha:
parallel
beta-beta
Future work
• Further refinement of binary contact scoring functions
- incorporate different contact types
- beta sheet vs. alpha helix
• Development of search procedures to explore contact map
space
• Other unrelated stuff
- proteomics
- gene expression and divergence
- physicochemical pattern recognition
Protein Threading
❑ Step 1: Build a contact map for the template structure
➢ contact map indicates interactions between cores, i.e. if any two residues
within the cores interact
Xu et al., JBCB, 2003
Protein Threading
❑ Step 3: Build a contact map for the query and template
Xu et al., JBCB, 2003
Protein Threading
❑ Step 4: define energy function and carry out minimization
Xu et al., JBCB, 2003