0% found this document useful (0 votes)
36 views12 pages

Molecular Docking in Drug Discovery

Uploaded by

Nour.Najim1900
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
36 views12 pages

Molecular Docking in Drug Discovery

Uploaded by

Nour.Najim1900
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

146 Current Computer-Aided Drug Design, 2011, 7, 146-157

Molecular Docking: A Powerful Approach for Structure-Based Drug


Discovery
Xuan-Yu Meng1,2, Hong-Xing Zhang*,1, Mihaly Mezei3 and Meng Cui*,2

1
State Key Laboratory of Theoretical and Computational Chemistry, Institute of Theoretical Chemistry, Jilin University,
Changchun, 130023, China
2
Department of Physiology and Biophysics, Virginia Commonwealth University, 1101 East Marshall Street, P.O. Box
980551, Richmond, VA 23298, USA
3
Department of Structural and Chemical Biology, Mount Sinai School of Medicine, Box 1677, New York, NY 10029,
USA

Abstract: Molecular docking has become an increasingly important tool for drug discovery. In this review, we present a
brief introduction of the available molecular docking methods, and their development and applications in drug discovery.
The relevant basic theories, including sampling algorithms and scoring functions, are summarized. The differences in and
performance of available docking software are also discussed. Flexible receptor molecular docking approaches, especially
those including backbone flexibility in receptors, are a challenge for available docking methods. A recently developed
Local Move Monte Carlo (LMMC) based approach is introduced as a potential solution to flexible receptor docking
problems. Three application examples of molecular docking approaches for drug discovery are provided.
Keywords: Conformational sampling, scoring, flexible protein-ligand docking, backbone flexibility, local move Monte Carlo
sampling.

INTRODUCTION important tool in pharmaceutical research. Various excellent


reviews on docking have been published in the past [5, 11-
The completion of the human genome project has 14], and many comparison studies were conducted to
resulted in an increasing number of new therapeutic targets evaluate the relative performance of the programs [15-18].
for drug discovery. At the same time, high-throughput
protein purification, crystallography and nuclear magnetic The molecular docking approach can be used to model
resonance spectroscopy techniques have been developed and the interaction between a small molecule and a protein at the
contributed to many structural details of proteins and atomic level, which allows us to characterize the behavior of
protein-ligand complexes. These advances allow the small molecules in the binding site of target proteins as well
computational strategies to permeate all aspects of drug as to elucidate fundamental biochemical processes [19]. The
discovery today [1-5], such as the virtual screening (VS) docking process involves two basic steps: prediction of the
techniques [6] for hit identification and methods for lead ligand conformation as well as its position and orientation
optimization. Compared with traditional experimental high- within these sites (usually referred to as pose) and
throughput screening (HTS), VS is a more direct and rational assessment of the binding affinity. These two steps are
drug discovery approach and has the advantage of low cost related to sampling methods and scoring schemes,
and effective screening [7-9]. VS can be classified into respectively, which will be discussed in the theory section.
ligand-based and structure-based methods. When a set of Knowing the location of the binding site before docking
active ligand molecules is known and little or no structural processes significantly increases the docking efficiency. In
information is available for targets, the ligand-based many cases, the binding site is indeed known before docking
methods, such as pharmacophore modeling and quantitative ligands into it. Also, one can obtain information about the
structure activity relationship (QSAR) methods can be sites by comparison of the target protein with a family of
employed. As to structure-based drug design, molecular proteins sharing a similar function or with proteins co-
docking is the most common method which has been widely crystallized with other ligands. In the absence of knowledge
used ever since the early 1980s [10]. Programs based on about the binding sites, cavity detection programs or online
different algorithms were developed to perform molecular servers e.g. GRID [20, 21], POCKET [22], SurfNet [23, 24],
docking studies, which have made docking an increasingly PASS [25] and MMC [26] can be utilized to identify putative
active sites within proteins. Docking without any assumption
about the binding site is called blind docking.
*Address correspondence to these authors at the State Key Laboratory of
Theoretical and Computational Chemistry, Institute of Theoretical The early elucidation for the ligand-receptor binding
Chemistry, Jilin University, Changchun, 130023, China; Tel: +86-431- mechanism is the lock-and-key theory proposed by Fischer
88498966; Fax: +86-431-88498966; E-mail: zhanghx@[Link] or [27], in which the ligand fits into the receptor like lock and
Department of Physiology and Biophysics, Virginia Commonwealth
University, 1101 East Marshall Street, P.O. Box 980551, Richmond, VA
key. The earliest reported docking methods [10] were based
23298, USA; Tel: (804) 827-9931; Fax: (804) 828-7382; on this theory and both the ligand and receptor were treated
E-mail: mcui@[Link] as rigid bodies accordingly. Then the “induced-fit” theory

1573-4099/11 $58.00+.00 © 2011 Bentham Science Publishers Ltd.


Molecular Docking Current Computer-Aided Drug Design, 2011, Vol. 7, No. 2 147

[28, 29] created by Koshland takes the lock-and-key theory a calculated for a match; new ligand conformations are
step further, stating that the active site of the protein is governed by the distance matrix between the pharmacophore
continually reshaped by interactions with the ligands as the and the corresponding ligand atoms. Chemical properties,
ligands interact with the protein. This theory suggests that like hydrogen-bond donors and acceptors, can be taken into
the ligand and receptor should be treated as flexible during account during the match. Matching algorithms have the
docking. Consequently, it could describe the binding events advantage of speed; thus they may be used for the
more accurately than the rigid treatment. Considering the enrichment of active compounds from large libraries [7].
limitation of computer resources, docking has been Matching algorithms for ligand docking are available in
performed with a flexible ligand and a rigid receptor for a DOCK [10], FLOG [46], LibDock [47] and SANDOCK [48]
long time, and remains the most popular method in use [7, programs.
30-35]. Recently many efforts have been made to deal with Incremental construction (IC) [30, 49, 50] methods put
the flexibility of the receptor [36-42], however, flexible
the ligand into an active site in a fragmental and incremental
receptor docking, especially backbone flexibility in
fashion. The ligand is divided into several fragments by
receptors, still presents a major challenge for available
breaking its rotatable bonds and then one of these fragments
docking methods. In our study, we propose a Local Move
is selected to dock into the active site first. This anchor is
Monte Carlo (LMMC) approach as a potential solution to
usually the largest fragment or the piece which may have
flexible receptor docking problems. significant functional role or interaction with protein. The
remaining fragments can be added incrementally. Different
THEORY OF DOCKING orientations are generated to fit in the active site, which
realizes the flexibility of the ligand. The incremental
Essentially, the aim of molecular docking is to give a
construction method has been used in DOCK 4.0 [51], FlexX
prediction of the ligand-receptor complex structure using
[30], Hammerhead [52], SLIDE [53] and eHiTS [54].
computation methods. Docking can be achieved through two
interrelated steps: first by sampling conformations of the In addition to IC, Multiple Copy Simultaneous Search
ligand in the active site of the protein; then ranking these (MCSS) [55, 56] and LUDI [57] are fragment-based
conformations via a scoring function. Ideally, sampling methods for the de novo design of ligands and modifications
algorithms should be able to reproduce the experimental of known ligands that may enhance their binding to the
binding mode and the scoring function should also rank it target protein. MCSS makes 1,000 to 5,000 copies of a
highest among all generated conformations. From these two functional group, which are randomly placed in the binding
perspectives, we give a brief overview of basic docking site of interest and subjected to simultaneous energy
theory. minimization and/or quenched molecular dynamics in the
forcefield of the protein. Copies only interact with the
Sampling Algorithms proteins and any interactions among the copies are omitted.
Consequently a set of energetically favorable binding sites
With six degrees of translational and rotational freedom and orientations for the functional group is identified based
as well as the conformational degrees of freedom of both the on the interaction energies. The binding site is mapped by
ligand and protein, there are a huge number of possible using different functional groups. New molecules which
binding modes between two molecules. Unfortunately, it perfectly match the binding site can be designed through the
would be too expensive to computationally generate all the linkage of those different functional groups.
possible conformations. Various sampling algorithms have LUDI focuses on the hydrogen bonds and hydrophobic
been developed and widely used in molecular docking contacts which could be formed between the ligand and
software (Table 1). protein. Its central concept are interaction sites, which are
Matching algorithms (MA) [43-45] based on molecular discrete positions in space suitable for forming hydrogen
shape map a ligand into an active site of a protein in terms of bonds or for filling a hydrophobic pocket [57]. A set of
shape features and chemical information. The protein and the interaction sites is generated either by searching the database
ligand are represented as pharmacophores. Each distance of or using the rules. The fragment is then fitted onto the
the pharmacophore within the protein and ligand is interaction sites and evaluated by distance criteria. The final

Table 1. Some Sampling Algorithms Discussed in this Paper

Algorithms Characteristic Ref.

Matching algorithms Geometry-based, suitable to VS and database enrichment for its high speed [43-45]
Incremental construction Fragment-based and docking incrementally [30, 49, 50]
MCSS fragment-based methods for the de novo design [55, 56]
LUDI fragment-based methods for the de novo design [57]
Monte Carlo Stochastic search [58, 59]
Genetic algorithms Stochastic search [31, 32, 64]
Molecular dynamics For further refinement after docking [68-70]
148 Current Computer-Aided Drug Design, 2011, Vol. 7, No. 2 Meng et al.

step is the connection of some or all of the fitted fragments Genetic algorithms have been used in AutoDock [31],
to a single molecule. GOLD [65], DIVALI [66] and DARWIN [67].
Stochastic methods search the conformational space by Molecular dynamics (MD) [68-70] is widely used as a
randomly modifying a ligand conformation or a population powerful simulation method in many fields of molecular
of ligands. Monte Carlo (MC) and genetic algorithms are modeling. In the context of docking, by moving each atom
two typical algorithms that belong to the class of stochastic separately in the field of the rest atoms, MD simulation
methods. represents the flexibility of both the ligand and protein more
effectively than other algorithms. However, the disadvantage
Monte Carlo (MC) [58, 59] methods generate poses of
of MD simulations is that they progress in very small steps
the ligand through bond rotation, rigid-body translation or
rotation. The conformation obtained by this transformation is and thus have difficulties in stepping over high energy
conformational barriers, which may lead to inadequate
tested with an energy-based selection criterion. If it passes
sampling. On the other hand, MD simulations are often
the criterion, it will be saved and further modified to
efficient at local optimization. Thus a current strategy is to
generate the next conformation. The iterations will proceed
use random search in order to identify the conformation of
until the pre-defined quantity of conformations is collected.
the ligand, followed by the further subtle MD simulations.
The main advantage of MC is that the change can be quite
large allowing the ligand to cross the energy barriers on the
potential energy surface, a point that isn’t achieved easily by Scoring Functions
molecular dynamics based simulation methods. Examples of
The purpose of the scoring function is to delineate the
applying the Monte Carlo methods include an earlier version
correct poses from incorrect poses, or binders from inactive
of AutoDock [60], ICM [61], QXP [62] and Affinity [63].
compounds in a reasonable computation time. However,
Genetic algorithms (GA) [31, 32, 64] form another class scoring functions involve estimating, rather than calculating
of well-known stochastic methods. The idea of the GA stems the binding affinity between the protein and ligand and
from Darwin’s theory of evolution. Degrees of freedom of through these functions, adopting various assumptions and
the ligand are encoded as binary strings called genes. These simplifications. Scoring functions can be divided in force-
genes make up the ‘chromosome’ which actually represents field-based, empirical and knowledge-based scoring
the pose of the ligand. Mutation and crossover are two kinds functions [5]. Table 2 shows some examples of scoring
of genetic operators in GA. Mutation makes random changes function formulae belonging to those three classes of scoring
to the genes; crossover exchanges genes between two functions respectively.
chromosomes. When the genetic operators affect the genes,
Classical force-field-based scoring functions [71-73]
the result is a new ligand structure. New structures will be
assess the binding energy by calculating the sum of the non-
assessed by scoring function, and the ones that survived (i.e.
bonded (electrostatics and van der Waals) interactions. The
exceeded a threshold) can be used for the next generation.

Table 2. Examples of Scoring Function Formulae

Scoring Function Formulae Ref.

A B  C D  qq ( r 2 )
V = Wvdw   12ij  6ij + Whbond  E (t ) 12ij  10ij + Welec  i j + Wsol  ( SiV j + S jVi )e ij
2 2
 r
i , j  ij r  r r  ( r ) r
ij i, j  ij ij i, j ij ij i, j
[31]
Extended force-field-based scoring function from AutoDock.
For two atoms i, j, the pair-wise atomic energy is evaluated by the sum of van der Waals, hydrogen bond, coulomb energy and desolvation. W are
weight factor to calibrate the empirical free energy.

G = G0 + Grot  N rot + Ghb  f (R,  )+ Gio  f (R,  )+ G  f (R,  )+ G


aro lipo  f (R ) *

neutral H bond ion init . aro int . lipo cont .

Empirical scoring function from FlexX.


G is the estimated free energy of binding; G0 is the regression constant; Grot , Ghb , Gio , Garo and Glipo are regression coefficients [30]

for each corresponding free energy term; f (R,  ) is scaling function penalizing deviations from the ideal geometry; N rot is the number of
free rotate bonds that are immobilized in the complex.

  ij (r )
PMF_ score =  A (r ) ij Aij (r ) = k BT ln  fVolj _ corr (r ) segij 
bulk 
kI 
r < rcut
ij
 off

Knowledge-based scoring functions PMF. [84]


k B is the Boltzmann constant; T is the absolute temperature; r is the atom pair distance. fVolj _ corr (r ) is the ligand volume correction factor;
 seg
ij
(r ) designates the radial distribution function of a protein atom of type i and a ligand atom of type j.
 bulk
ij
Molecular Docking Current Computer-Aided Drug Design, 2011, Vol. 7, No. 2 149

electrostatic terms are calculated by a Coulombic selection of proteins for successful structure determination
formulation. Since such point charge calculations have thus the obtained parameters may not be suitable for
problems in modeling the protein’s real environment a widespread use, especially with interactions involving metals
distance-dependent dielectric function is generally used to or halogens. PMF [84], DrugScore [90], SMoG [91] and
modulate the contribution of charge-charge interactions. The Bleep [85] are examples of knowledge-based functions
van der Waals terms are described by a Lennard-Jones which differ mainly in the size of training sets, the form of
potential function. Adopting different parameter sets for the the energy function, the definition of atom types, distance
Lennard-Jones potential can vary the “hardness” of the cutoff or other parameters.
potential which controls how close a contact between protein
Consensus scoring [92] is a recent strategy that combines
and ligand atoms can be acceptable. Force-field-based
several different scores to assess the docking conformation.
scoring functions also have the problem of slow A pose of ligand or a potential binder could be accepted
computational speed. Thus cut-off distance is used to handle
when it scores well under a number of different scoring
the non-bonded interactions. This also results in decreasing
schemes. Consensus scoring usually substantially improves
the accuracy of long-range effects involved in binding.
enrichments (i.e. the percentage of strong binder among the
Extensions of force-field-based scoring functions high-scoring ligands) in virtual screening, and improves the
consider the hydrogen bonds, solvations and entropy prediction of bound conformations and poses [93]. However,
contributions. Software, such as DOCK [10, 50, 51, 74], the prediction of binding energies might still be inaccurate.
GOLD [65] and AutoDock [31], offer users such functions. Also, the usefulness of consensus scoring diminishes when
They have some differences in the treatment of hydrogen terms in different scoring functions are significantly
bonds, the form of the energy function etc. Furthermore, the correlated [5, 93]. CScore [94] is an example of which
results of docking with force-field-based functions can be combines DOCK, ChemScore, PMF, GOLD, and FlexX
further refined with other techniques, such as linear scoring functions.
interaction energy [75] and free-energy perturbation methods
Typical scoring functions face the problem of affinity
(FEP) [71, 76] to improve the accuracy in predicting binding
prediction partly because of the limited treatment of
energies.
solvation effect. One of the ways to solve this problem is
In empirical scoring functions [77-81], binding energy physics-based scoring e.g. MM-PB/SA and MM-GB/SA
decomposes into several energy components, such as (MM stands for molecular mechanics, PB and GB for
hydrogen bond, ionic interaction, hydrophobic effect and Poisson-Boltzmann and Generalized Born, respectively, SA
binding entropy. Each component is multiplied by a for solvent-accessible surface area), which is involved in
coefficient and then summed up to give a final score. rescoring or lead optimization to improve the accuracy of
Coefficients are obtained from regression analysis fitted to a binding affinity prediction. Promising results were obtained
test set of ligand-protein complexes with known binding using MM-PB/SA [95, 96] or MM-GB/SA [97] in some
affinities. studies. However, recently Guimarães and Mathiowetz
reported that the GB/SA model poorly estimated protein
Empirical scoring functions have relatively simple
desolvation on certain systems, while incorporating
energy terms to evaluate. However, it is unclear as to how
WaterMap into the MM-GB/SA method instead of GB/SA
well they are suited for ligand-protein complexes beyond the
protein desolvation gave the best ranking result [98]. Singh
training set. Additionally, each term in empirical scoring
and Warshel compared several methods for evaluating the
functions may be treated in a different manner by different
software, and the numbers of the terms included are also affinity of protein-ligand complexes and suggested that
PDLD/S-LRA/ (protein dipoles Langevin dipoles linear
different. LUDI [57], PLP [78, 79, 82], ChemScore [83] are
response approximation) appears to offer an appealing
examples derived from empirical scoring functions
option for the final stages of massive VS and in contrast,
Knowledge-based scoring functions [84-89] use PB/SA appears to provide erroneous estimates of the
statistical analysis of ligand-protein complexes crystal absolute binding energies because of its incorrect estimation
structures to obtain the interatomic contact frequencies of entropies and the problematic treatment of electrostatic
and/or distances between the ligand and protein. They are energies [99].
based on the assumption that the more favorable an
interaction is, the greater the frequency of occurrence will DOCKING METHODOLOGIES
be. These frequency distributions are further converted into
pairwise atom-type potentials. The score is calculated by Rigid Ligand and Rigid Receptor Docking
favoring preferred contacts and penalizing repulsive
interactions between each atom in the ligand and protein When the ligand and receptor are both treated as rigid
within a given cutoff. bodies, the search space is very limited, considering only
three translational and three rotational degrees of freedom. In
The appeal of knowledge-based functions is this case, ligand flexibility could be addressed by using a
computational simplicity, which can be exploited to screen pre-computed a set of ligand conformations, or by allowing
large compound databases. They can also model some for a degree of atom-atom overlap between the protein and
uncommon interactions like sulphur-aromatic or cation-, ligand. The early versions of DOCK [10, 50, 51, 74], FLOG
which are often poorly handled in empirical approaches. [46] and some protein-protein docking programs, such as
However, they are still faced with the problem that some FTDOCK [100], adopted such a method that kept the ligand
interactions are underrepresented in the limited training sets and receptor rigid during the process of the docking.
of crystal structures as well as by the bias inherent in the
150 Current Computer-Aided Drug Design, 2011, Vol. 7, No. 2 Meng et al.

DOCK is the first automated procedure for docking a and metal and aromatic ring interactions between the ligand
molecule into a receptor site and is being continuously and protein. Then the remaining components are
developed. It characterizes the ligand and receptor as sets of incrementally built-up in accordance with a set of predefined
spheres which could be overlaid by means of a clique rotatable torsion angles to account for ligand flexibility. The
detection procedure [101]. Geometrical and chemical FlexX scoring function is based on Böhm’s work [105]. Its
matching algorithms are used, and the ligand-receptor current version includes terms of electrostatic interactions,
complexes can be scored by accounting for steric fit, directional hydrogen bonds, rotational entropy, and aromatic
chemical complementation or pharmacophore similarity. and lipophilic interactions. The interactions between
Within its improved versions, incremental construction functional groups are also taken into account through
method and exhaustive search are added to consider the assigning the type and geometry for groups.
ligand flexibility. The exhaustive search randomly generates
a user-defined number of conformers as a multiple of the Flexible Ligand and Flexible Receptor Docking
number of rotatable bonds in the ligand. With respect to
scoring, the latest version DOCK 6.4 has included both an The intrinsic mobility of proteins has been proven to be
AMBER-derived force-field scoring with implicit solvent closely related to ligand binding behavior and it has been
[102] and GB/SA, PB/SA solvation scoring [97, 103]. reviewed by Teague [106]. Incorporating the receptor
FLOG generates ligand conformations on the basis of flexibility is significant challenge in the field of docking.
Ideally, using MD simulations could model all the degrees of
distance geometry and uses a clique-finding algorithm to
freedom in the ligand-receptor complex. But MD has the
calculate the sets of distances. Up to 25 explicit
problem of inadequate sampling that we mentioned earlier.
conformations of the ligand could be used to dock for some
Another hurdle is its high computational expense, which
flexibility. FLOG allows users to define essential points
prevents this method from being used in the screening of
which must be paired with a ligand atom. This approach is
useful if an important interaction is already known before large chemical database.
docking. Conformations are scored with a function In addition to the historic induced fit several theoretical
considering van der Waals, electrostatics, hydrogen bonding models, conformer selection and conformational induction,
and hydrophobic interactions. have been proposed to illustrate the flexible ligand-protein
binding process. According to the definition given by
Flexible Ligand and Rigid Receptor Docking Teague [106], conformer selection refers to a process when a
ligand selectively binds to a favorable conformation from a
For systems whose behavior follows the induced fit number of protein conformations; conformational induction
paradigm [28, 29], it is of vital importance to consider the describes a process in which the ligand converts the protein
flexibilities of both the ligand and receptor since in that case into a conformation that it would not spontaneously adopt in
both the ligand and receptor change their conformations to its unbound state. In some cases, this conformational
form a minimum energy perfect-fit complex. However, the conversion can be likened to a partial refolding of the
cost is very high when the receptor is also flexible. Thus the protein.
common approach, also a trade-off between accuracy and
Various methods are currently available to implement the
computational time, is treating the ligand as flexible while
receptor flexibility (Table 3). The simplest one is so-called
the receptor is kept rigid during docking. Almost all the
“soft-docking” [37, 107, 108], decreases the van der Waals
docking programs have adopted this methodology, such as
repulsion energy term in the scoring function to allow for a
AutoDock [31], FlexX [30]. degree of atom-atom overlap between the receptor and
AutoDock 3.0 incorporates Monte Carlo simulated ligand. For example, the LJ 8-4 potential in GOLD and
annealing, evolutionary, genetic and Lamarckian genetic smooth potential in AutoDock 3.0 belong to this class. This
algorithm methods to model the ligand flexibility while method may not include adequate flexibility. Nevertheless, it
keeping the receptor rigid. The scoring function is based on has the advantage of computational efficiency as the receptor
the AMBER force field, including van der Waals, hydrogen coordinates are fixed, simply by adjusting van der Waals
bonding, electrostatic interactions, conformational entropy parameters.
and desolvation terms. Each term is weighted using an
Utilizing rotamer libraries [109, 110] is another approach
empirical scaling factor obtained from experimental data.
to modeling receptor flexibility. Rotamer libraries include a
AutoDock 4.0 is able to model receptor flexibility by
set of side-chain conformations which are usually
allowing side-chains to move. Additionally, interaction of
determined from statistical analysis of structural
protein-protein docking could be evaluated in this version of experimental data. The advantage of using rotamers is the
AutoDock. AutoDock Vina was recently released as the
relative speed in sampling, and the avoiding of minimization
latest version for molecular docking and virtual screening
barriers. ICM (Internal Coordinates Mechanics) [61] is a
[104]. By redocking the 190 receptor-ligand complexes that
program using rotamer libraries with the biased probability
had been used as a training set for the AutoDock 4,
methodology [111], coupled with Monte Carlo search of the
AutoDock Vina simultaneously showed approximately a two
ligand conformation.
orders exponential improvement of magnitude in speed and a
significantly better accuracy of the binding mode prediction. AutoDock 4 [112] adopts a simultaneous sample method
to deal with side chain flexibility. Several side chains of the
FlexX uses an incremental construction algorithm to
receptor can be selected by users and simultaneously
sample ligand conformations. The base fragment is first
sampled with a ligand using the same methods. Other
docked into the active site by matching hydrogen bond pairs
portions of the receptor are treated rigidly with a grid energy
Molecular Docking Current Computer-Aided Drug Design, 2011, Vol. 7, No. 2 151

Table 3. Some Basic Methods for Including Receptor Flexibility

Method Description Advantage Disadvantage Program

Change vdW to allow for


Computational efficiency. Inadequate flexibility.
overlap
Soft potential Easy to implement and use Describe flexibility in an implicit, AutoDock [31]
between receptor and ligand
combined with other methods. rude and non-quantitative way.
atoms
Relative computational Strong dependence on the database
Search side chain library to efficiency. used.
Rotamer library ICM [61]
obtain possible conformations
Avoid minimization barriers. No backbone flexibility.
Relative computational
Sample both side chain and Only selected side chains are
Receptor side chain efficiency. GOLD [65]
ligand conformations involved.
flexibility Model the effect that ligand AutoDock 4 [112]
simultaneously using GA No backbone flexibility.
make on binding site residues.
Docking ligand to a series of Expensive computational cost.
Ensemble of protein receptor structures which Include full and explicit DOCK [113]
conformations represent different flexibility. Limited by protein conformations
FlexE [38]
conformational states. used in sampling.

map during sampling. Grid energy map introduced by mean-field theory is applied to model induced-fit
Goodford [20] is used to store energy information of the complementarities between the ligand and protein.
receptor and simplify interaction energy calculation between
Methods mentioned above either include only side chain
ligand and receptor.
flexibility or full flexibility of the receptor. We have known
Still another way to deal with the protein flexibility is to that loops forming active sites play an important role in
use an ensemble of protein conformations, which ligand binding. In some cases the loop may undergo
corresponds to the theory of conformer selection [113, 114]. dramatic conformational change whereas in other portions of
A ligand is separately docked into a set of rigid protein the receptor there is little change upon ligand binding. For
conformations rather than a single one, and the results are this situation, side chain flexibility methods fail to sample
merged depending on the method of choice [115]. This the correct protein conformation and full flexibility seems to
method was originally implemented in DOCK, which be a computational waste. Fig. (1) shows superimposed
generates an average potential energy grid of the ensemble crystal structures of triosephosphate isomerase as an
[113] and is extended in many programs in different ways. example. The active site of triosephosphate isomerase has an
For example, FlexE [38] collects multiple crystal structures 11-residue loop which moves 7Å upon ligand binding [116].
of a certain protein, merging the similar parts while marking However, the rest of the enzyme has no movement in
the dissimilar areas as different alternatives. During the comparison to their apo and holo structures. Several enzyme
incremental construction of a ligand discrete protein families also involve loop rearrangement within the active
conformations are sampled in a combinatorial fashion. The site responsible for ligand binding, such as Bromodomain, an
highest scoring protein structure is selected based on a extensive family related to acetyl-lysine binding, or
comparison between the ligand and each alternative. Dihydrofolate reductase, responsible for the maintenance of
the cellular pools of tetrahydrofolate, as well as other kinds
Hybrid method is another practical strategy to model
of kinases [117, 118]. In the next section, we present the
receptor flexibility. One example is Glide [33], a very
Local Move Monte Carlo (LMMC) loop sampling method, a
popular program in the field of docking. Glide designs a
series of hierarchical filters to search the possible poses and new approach which focuses on sampling ligand
conformation within loop-containing active sites.
orientations of the ligand within the binding site of the
receptor. Ligand flexibility is handled by an exhaustive
search of the ligand torsion angle space. Initial ligand Local Move Monte Carlo Sampling for Flexible Receptor
conformations are selected based on torsion energies and Docking
docked into receptor binding sites with soft potentials. Then
a rotamer exploration is used to further model receptor Local move (also referred to as ‘window move’) starts
with changing one torsion angle (called the driver torsion)
flexibility [36]. IFREDA [115] utilizes a hybrid method that
followed by the adjustment of the six subsequent torsions to
combines soft potential and multiple receptor conformations,
allow the rest of the chain to remain in its original position
accounting for receptor flexibility. Other programs, like
while preserving all bond lengths and bond angles (Fig. 2).
QXP [62] and Affinity [63], perform a Monte Carlo search
The pioneering work on local move was done by Go and
of ligand conformations followed by a minimization step.
During minimization, the user-defined parts of the protein Scheraga [119], who developed a solution for the system of
equations defining the values of the six torsion angles that
are allowed to move in order to avoid atom clashes between
preserve the backbone bond lengths and angles. Hoffmann
the ligand and receptor. SLIDE [53] is designed to
and Knapp first applied the local move method in a MC
incorporate flexibility with the ability to remove clashes by
simulation of polyalanine folding that included a suitable
directed, single bond rotation of either the ligand or the side
Jacobian [120], required for maintaining detailed balance.
chains of the protein. An optimization approach based on the
They demonstrated that this method samples the
152 Current Computer-Aided Drug Design, 2011, Vol. 7, No. 2 Meng et al.

conformational space more efficiently than single move methods. The results show that this approach can reproduce
[121]. The method has been further tested on proline- the experimental results with root mean square deviation
containing peptides [122], proteins and nucleic acids [123]. (RMSD) within 1.8 Å for the all the test cases [125]. Fig. (3)
Mezei introduced the ‘reverse proximity criterion’ for shows the loop structures of 2act (198-205) sampled by the
filtering all possible loop closure solutions to select the most LMMC method. This LMMC loop prediction approach
structurally conservative one and tested it on a solvated lipid could be useful for flexible receptor docking. In our future
bilayer [124]. studies, we will develop our LMMC based molecular
docking approach, which samples not only the side chains
but also the backbone loops in the binding site of proteins
and flexible ligands as well. A flowchart of the LMMC
based molecular docking approach is given in Fig. (4).

Fig. (1). Superimposed apo- (without ligand, in black) and holo- Fig. (3). Loop Structure of 2act (198-205) produced by the local
(with ligand in gray) crystal structures of triosephosphate move MC method at 5000K and followed by clustering to generate
isomerase. PDB code 1YPI and 2YPI, respectively [116]. The 11 100 representative conformations. Black stick represents the crystal
residue-loop composed of binding site is the only region that has loop structure, and gray wires represent the 100 representative loop
large motion upon ligand binding (in circle). conformations.

APPLICATION EXAMPLES OF MOLECULAR DOC-


KING FOR DRUG DISCOVERY
Molecular docking has been the most widely employed
technique. Though the main application lies in structure-
based virtual screening for identification of new active
compounds towards a particular target protein, in which it
has produced a number of success stories [126], it is actually
not a stand-alone technique but is normally embedded in a
Fig. (2). Local move of a lipid tail. Six subsequent torsions change
workflow of different in silico as well as experimental
while keeping the rest of the chain to remain in its original position.
techniques [127]. Several research groups focus on
We have developed an improved local move Monte evaluating of the performance of various docking programs
Carlo (LMMC) loop sampling approach for loop predictions. or on making improvements to the scoring functions when
The method generates loop conformations based on simple experimental testing has already been done. Such efforts
moves of the torsion angles of side chains and local moves could give meaningful guidance to choose the methodology
of the backbones of loops. To reduce the computational costs for a particular target system. Docking, combined with other
for energy evaluations, we developed a grid-based force field computational techniques and experimental data, also could
to represent the protein environment and solvation effect. be involved in analyzing drug metabolism to obtain some
Simulated annealing has been used to enhance the efficiency useful information from the cytochrome P450 system [128-
of the LMMC loop sampling and identify low-energy loop 130], for example. In the following, three examples of
conformations. The prediction quality was evaluated on a set successful applications of docking are presented.
of protein loops with a known crystal structure that has been
DNA gyrase is a bacterial enzyme that introduces
previously used by others to test different loop prediction negative supercoils into bacterial DNA and unwinds of
Molecular Docking Current Computer-Aided Drug Design, 2011, Vol. 7, No. 2 153

Fig. (4). Flowchart of local move Monte Carlo (LMMC) loop sampling approach for protein-ligand docking. Abbreviate: MC: Monte Carlo;
ESC: Exponential Coolong Schedule; LCS: Linear Cooling Schedule.

DNA, thus being studied as antibacterial target. HTS failed significantly lower than the value of 79% observed with both
to find novel inhibitors of DNA gyrase. Boehm et al. used de scoring functions for the full GOLD validation set.
novo design for this enzyme and successfully obtained Additionally, it is apparent from the data that the search
several new inhibitors [131]. Firstly, 3D complex structures algorithm was very unlikely to be responsible for the failure
of DNA gyrase with known inhibitors, ciprofloxacin and in docking. Further research indicated that re-
novobiocin, were carefully analyzed to get a common parameterization of metal-acceptor interactions and
binding pattern, in which both inhibitors donate one lipophilicity of planar nitrogen atoms in the scoring
hydrogen bond to Asp73 and accept one hydrogen bond functions resulted in a significant increase in the percentage
from a conserved water molecule. In addition, some of successful docking poses against the heme binding
lipophilic fragments should be included in the molecule to proteins (Chemscore 73%, Goldscore 65%), which might be
have lipophilic interaction with the receptor. Based on this useful in docking applications on P450 enzymes and other
information, LUDI and CATALYST were employed to heme-binding proteins.
search the Available Chemicals Directory (ACD) and a part
Concerning VS and HTS, comparative research has been
of the Roche compound inventory (RIC), respectively, and done by Doman et al. [133]. Both VS and HTS were applied
collected about 600 compounds. Close analogs of these
to screen the inhibitors of the protein tyrosine phosphatase-
compounds were also considered, thus in total 3000
1B (PTP-1B). For the HTS a library of approximately
compounds were further tested using biased screening.
400,000 compounds from a corporate collection were
Consequently 150 hits were selected and clustered into 14
screened. Some 85 compounds were found with IC50 values
classes of which 7 classes were proven to be the true and
less than 100 μM, corresponding to a hit rate of 0.021%. And
novel inhibitors. Subsequent hit optimization relied strongly the most active had an IC50 value of 4.2 μM. For VS,
on the knowledge of 3D structures of the binding site and
235,000 commercially available molecules were docked into
eventually generated a series of highly potent DNA gyrase
the crystal structure of PTP-1B (PDB code 1pty) using the
inhibitors.
Northwestern University version [134-137] of DOCK3.5
Another example is focused on the validation of docking [102, 138]. After docking, the top-scoring 1000 molecules
and scoring applied in cytochromes P450 and other heme- (500 for the ACD and 500 for the combined BioSpecs and
containing proteins [132]. Docking against heme-containing Maybridge databases) were considered for further
complexes appears to be difficult because certain ligands evaluation. A total of 889 molecules were actually available,
coordinate directly to the heme iron atom and the precise and after visual inspection 365 compounds were chosen for
energetics of this contact for different chelating groups needs testing. Of these, 127 molecules were found to be active with
to be properly balanced with other energetic terms, and in the IC50 <100 μM, corresponding to a hit rate of 34.8%.
case of the P450s, the environment above the heme group is Structure-based docking therefore enriched the hit rate by
very hydrophobic compared to other enzymes and some 1700-fold over random screening. Another point that should
scoring functions and docking methods perform poorly on be noted is that the hits from VS and HTS are very different
interactions driven entirely by lipophilic contacts. In this from each other, which implies that combination of VS and
study, 45 complexes from the PDB database comprising HTS may be more helpful for lead discovery.
heme-containing proteins and ligands were selected. The
native ligands were removed and then docked into the CONCLUDING REMARKS
defined active cavities using the GOLD [65] software which
employs genetic algorithms to generate ligand Receptor flexibility, especially backbone flexibility and
conformations. The scoring functions used to rank the movement of several key secondary elements of the receptor
docking poses were Goldscore [32] and Chemscore [65]. involving ligand binding and the catalyst, is still a major
The results show that the success rates are 64% and 57% for hurdle in docking studies. Some methods to deal with side
Chemscore and Goldscore respectively, which is chain flexibility have been proven effective and adequate in
154 Current Computer-Aided Drug Design, 2011, Vol. 7, No. 2 Meng et al.

certain cases. With respect to global flexibility, an ensemble [13] Kontoyianni, M.; Madhav, P.; Suchanek, E.; Seibel, W. Theoretical
of proteins is a popular solution which accords with the and practical considerations in virtual screening: a beaten field?
Curr. Med. Chem., 2008, 15, 107-116.
viewpoint of conformer selection. It requires an efficient [14] Brooijmans, N.; Kuntz, I.D. Molecular recognition and docking
way to obtain and select reliable protein structures used for algorithms. Annu. Rev. Biophys. Biomol. Struct., 2003, 32, 335-
docking, which means structures that the ligand can fit in 373.
should be included in the ensembles. Besides, computational [15] ten Brink, T.; Exner, T.E. Influence of protonation, tautomeric, and
stereoisomeric states on protein-ligand docking results. J. Chem.
cost is another limitation for this method. LMMC could be Inf. Model., 2009, 49, 1535-1546.
an appropriate method for sampling a ligand within loop- [16] Cross, J.B.; Thompson, D.C.; Rai, B.K.; Baber, J.C.; Fan, K.Y.;
containing active sites since loop tends to be more flexible Hu, Y.; Humblet, C. Comparison of several molecular docking
and hard to model using existing approaches especially due programs: pose prediction and virtual screening accuracy. J. Chem.
to their possibly dramatic movements. Another advantage is Inf. Model., 2009, 49, 1455-1474.
[17] Li, X.; Li, Y.; Cheng, T.; Liu, Z.; Wang, R. Evaluation of the
the adjustment of the extent of flexibility. Either the side performance of four molecular docking programs on a diverse set
chain or full movement of the loop can be directly controlled of protein-ligand complexes. J. Comput. Chem., 2010, 31, 2109-
by users. 2125.
[18] Plewczynski, D.; Lazniewski, M.; Augustyniak, R.; Ginalski, K.
Scoring function is a fundamental component worth Can we trust docking results? Evaluation of seven commonly used
being further improved upon in docking. Successful programs on PDBbind database. J. Comput. Chem., 2010
application examples show that computational approaches doi: 10.1002/jcc.21643.
[19] McConkey, B.J.; Sobolev, V.; Edelman, M. The performance of
have the power to screen hits from a huge database and current methods in ligand-protein docking. Curr. Sci., 2002, 83,
design novel small molecules. However, the realistic 845-855.
interactions between small molecules and receptors still rely [20] Goodford, P.J. A computational procedure for determining
on experimental technology. Accurate as well as low energetically favorable binding sites on biologically important
computational cost scoring functions may bring docking macromolecules. J. Med. Chem., 1985, 28, 849-857.
[21] Kastenholz, M.A.; Pastor, M.; Cruciani, G.; Haaksma, E.E.; Fox, T.
application to a new stage. GRID/CPCA: a new computational tool to design selective ligands.
J. Med. Chem., 2000, 43, 3033-3044.
ACKNOWLEDGEMENTS [22] Levitt, D.G.; Banaszak, L.J. POCKET: a computer graphics
method for identifying and displaying protein cavities and their
surrounding amino acids. J. Mol. Graph., 1992, 10, 229-234.
We gratefully acknowledge financial support from [23] Laskowski, R.A. SURFNET: a program for visualizing molecular
National Institute Health (DC008996 and S10RR027411 to surfaces, cavities, and intermolecular interactions. J. Mol. Graph.,
M.C.). The computations were partially supported by grants 1995, 13, 323-330, 307-328.
of the National Center for Supercomputing Applications [24] Glaser, F.; Morris, R.J.; Najmanovich, R.J.; Laskowski, R.A.;
(MCB070095T and MCB090019 to M.C.). Thornton, J.M. A method for localizing ligand binding pockets in
protein structures. Proteins, 2006, 62, 479-488.
[25] Brady, G.P. Jr.; Stouten, P.F. Fast prediction and visualization of
REFERENCES protein binding pockets with PASS. J. Comput. Aided Mol. Des.,
2000, 14, 383-401.
[1] Jorgensen, W.L. The many roles of computation in drug discovery. [26] Mezei, M. A new method for mapping macromolecular
Science, 2004, 303, 1813-1818. topography. J. Mol. Graph. Model., 2003, 21, 463-472.
[2] Bajorath, J. Integration of virtual and high-throughput screening. [27] Fischer, E. Einfluss der configuration auf die wirkung derenzyme.
Nat. Rev. Drug Discov., 2002, 1, 882-894. Ber. Dt. Chem. Ges., 1894, 27, 2985-2993.
[3] Walters, W.P.; Stahl, M.T.; Murcko, M.A. Virtual screening - an [28] Koshland, D.E. Jr. Correlation of structure and function in enzyme
overview. Drug Discov. Today, 1998, 3, 160-178 action. Science, 1963, 142, 1533-1541.
[4] Langer, T.; Hoffmann, R.D. Virtual screening: an effective tool for [29] Hammes, G.G. Multiple conformational changes in enzyme
lead structure discovery? Curr. Pharm. Des., 2001, 7, 509-527. catalysis. Biochemistry, 2002, 41, 8221-8228.
[5] Kitchen, D.B.; Decornez, H.; Furr, J.R.; Bajorath, J. Docking and [30] Rarey, M.; Kramer, B.; Lengauer, T.; Klebe, G. A fast flexible
scoring in virtual screening for drug discovery: methods and docking method using an incremental construction algorithm. J.
applications. Nat. Rev. Drug Discov., 2004, 3, 935-949. Mol. Biol., 1996, 261, 470-489.
[6] Gohlke, H.; Klebe, G. Approaches to the description and prediction [31] Morris, G.M.; Goodsell, D.S.; Halliday, R.S.; Huey, R.; Hart,
of the binding affinity of small-molecule ligands to W.E.; Belew, R.K.; Olson, A.J. Automated docking using a
macromolecular receptors. Angew. Chem. Int. Ed. Engl., 2002, 41, Lamarckian genetic algorithm and an empirical binding free energy
2644-2676. function. J. Comput. Chem., 1998, 19, 1639-1662.
[7] Moitessier, N.; Englebienne, P.; Lee, D.; Lawandi, J.; Corbeil, C.R. [32] Jones, G.; Willett, P.; Glen, R.C.; Leach, A.R.; Taylor, R.
Towards the development of universal, fast and highly accurate Development and validation of a genetic algorithm for flexible
docking/scoring methods: a long way to go. Br. J. Pharmacol., docking. J. Mol. Biol., 1997, 267, 727-748.
2008, 153(Suppl 1), S7-26. [33] Friesner, R.A.; Banks, J.L.; Murphy, R.B.; Halgren, T.A.; Klicic,
[8] Shoichet, B.K.; McGovern, S.L.; Wei, B.; Irwin, J.J. In: Hits, leads J.J.; Mainz, D.T.; Repasky, M.P.; Knoll, E.H.; Shelley, M.; Perry,
and artifacts from virtual and high throughput screening., J.K.; Shaw, D.E.; Francis, P.; Shenkin, P.S. Glide: a new approach
Molecular Informatics: Confronting Complexity, 2002. for rapid, accurate docking and scoring. 1. Method and assessment
[9] Bailey, D.; Brown, D. High-throughput chemistry and structure- of docking accuracy. J. Med. Chem., 2004, 47, 1739-1749.
based design: survival of the smartest. Drug Discov. Today, 2001, [34] McGann, M.R.; Almond, H.R.; Nicholls, A.; Grant, J.A.; Brown,
6, 57-59. F.K. Gaussian docking functions. Biopolymers, 2003, 68, 76-90.
[10] Kuntz, I.D.; Blaney, J.M.; Oatley, S.J.; Langridge, R.; Ferrin, T.E. [35] Perola, E.; Walters, W.P.; Charifson, P.S. A detailed comparison of
A geometric approach to macromolecule-ligand interactions. J. current docking and scoring methods on systems of pharmaceutical
Mol. Biol., 1982, 161, 269-288. relevance. Proteins, 2004, 56, 235-249.
[11] Halperin, I.; Ma, B.; Wolfson, H.; Nussinov, R. Principles of [36] Sherman, W.; Day, T.; Jacobson, M.P.; Friesner, R.A.; Farid, R.
docking: an overview of search algorithms and a guide to scoring Novel procedure for modeling ligand/receptor induced fit effects. J.
functions. Proteins, 2002, 47, 409-443. Med. Chem., 2006, 49, 534-553.
[12] Coupez, B.; Lewis, R.A. Docking and scoring--theoretically easy, [37] Jiang, F.; Kim, S.H. “Soft docking”: matching of molecular surface
practically impossible? Curr. Med. Chem., 2006, 13, 2995-3003. cubes. J. Mol. Biol., 1991, 219, 79-102.
Molecular Docking Current Computer-Aided Drug Design, 2011, Vol. 7, No. 2 155

[38] Claussen, H.; Buning, C.; Rarey, M.; Lengauer, T. FlexE: efficient prediction from the distorted native conformation. J. Comput.
molecular docking considering protein structure variations. J. Mol. Chem., 1994, 15, 488-506.
Biol., 2001, 308, 377-395. [62] McMartin, C.; Bohacek, R.S. QXP: powerful, rapid computer
[39] Alonso, H.; Bliznyuk, A.A.; Gready, J.E. Combining docking and algorithms for structure-based drug design. J. Comput. Aided Mol.
molecular dynamic simulations in drug design. Med. Res. Rev., Des., 1997, 11, 333-344.
2006, 26, 531-568. [63] Accelrys Inc., San Diego, CA, USA.
[40] Sander, T.; Liljefors, T.; Balle, T. Prediction of the receptor [64] Oshiro, C.M.; Kuntz, I.D.; Dixon, J.S. Flexible ligand docking
conformation for iGluR2 agonist binding: QM/MM docking to an using a genetic algorithm. J. Comput. Aided Mol. Des., 1995, 9,
extensive conformational ensemble generated using normal mode 113-130.
analysis. J. Mol. Graph. Model., 2008, 26, 1259-1268. [65] Verdonk, M.L.; Cole, J.C.; Hartshorn, M.J.; Murray, C.W.; Taylor,
[41] Subramanian, J.; Sharma, S.; B-Rao, C. A novel computational R.D. Improved protein-ligand docking using GOLD. Proteins,
analysis of ligand-induced conformational changes in the ATP 2003, 52, 609-623.
binding sites of cyclin dependent kinases. J. Med. Chem., 2006, 49, [66] Clark, K.P.; Ajay, Flexible ligand docking without parameter
5434-5441. adjustment across four ligand- receptor complexes. J. Comput.
[42] Subramanian, J.; Sharma, S.; B-Rao, C. Modeling and selection of Chem., 1995, 16, 1210-1226.
flexible proteins for structure-based drug design: backbone and side [67] Taylor, J.S.; Burnett, R.M. DARWIN: a program for docking
chain movements in p38 MAPK. Chem. Med. Chem., 2008, 3, 336- flexible molecules. Proteins, 2000, 41, 173-191.
344. [68] Cornell, W.D.; Cieplak, P.; Bayly, C.I.; Gould, I.R.; Merz, K.M.;
[43] Brint, A.T.; Willett, P. Algorithms for the identification of three- Ferguson, D.M.; Spellmeyer, D.C.; Fox, T.; Caldwell, J.W.;
dimensional maximal common substructures. J. Chem. Inf. Kollman, P.A. A second generation force field for the simulation of
Comput. Sci., 1987, 27, 152-158. proteins, nucleic acids, and organic molecules. J. Am. Chem. Soc.,
[44] Fischer, D.; Norel, R.; Wolfson, H.; Nussinov, R. Surface motifs by 1995, 117, 5179-5197.
a computer vision technique: searches, detection, and implications [69] Weiner, S.J.; Kollman, P.A.; Case, D.A.; Singh, U.C.; Ghio, C.;
for protein-ligand recognition. Proteins, 1993, 16, 278-292. Alagona, G.; Profeta, S. Jr.; Weiner, P. New force field for
[45] Norel, R.; Fischer, D.; Wolfson, H.J.; Nussinov, R. Molecular molecular mechanical simulation of nucleic acids and proteins. J.
surface recognition by a computer vision-based technique. Protein Am. Chem. Soc., 1984, 106, 765-784.
Eng., 1994, 7, 39-46. [70] Brooks, B.R.; Bruccoleri, R.E.; Olafson, B.D.; States, D.J.;
[46] Miller, M.D.; Kearsley, S.K.; Underwood, D.J.; Sheridan, R.P. Swaminathan, S.; Karplus, M. CHARMM: A program for
FLOG: a system to select ‘quasi-flexible’ ligands complementary macromolecular energy, minimization, and dynamics calculations.
to a receptor of known three-dimensional structure. J. Comput. J. Comput. Chem., 1983, 4, 187-217.
Aided Mol. Des., 1994, 8, 153-174. [71] Kollman, P.A. Free energy calculations: Applications to chemical
[47] Diller, D.J.; Merz, K.M., Jr. High throughput docking for library and biochemical phenomena. Chem. Rev., 1993, 93 2395-2417.
design and library prioritization. Proteins, 2001, 43, 113-124. [72] Aqvist, J.; Luzhkov, V.B.; Brandsdal, B.O. Ligand binding
[48] Burkhard, P.; Taylor, P.; Walkinshaw, M.D. An example of a affinities from MD simulations. Acc. Chem. Res., 2002, 35, 358-
protein ligand found by database mining: description of the 365.
docking method and its verification by a 2.3 A X-ray structure of a [73] Carlson, H.A.; Jorgensen, W.L. An extended linear response
thrombin-ligand complex. J. Mol. Biol., 1998, 277, 449-466. method for determining free energies of hydration. J. Phys. Chem.,
[49] DesJarlais, R.L.; Sheridan, R.P.; Dixon, J.S.; Kuntz, I.D.; 1995 99, 10667-10673.
Venkataraghavan, R. Docking flexible ligands to macromolecular [74] Shoichet, B.K.; Stroud, R.M.; Santi, D.V.; Kuntz, I.D.; Perry, K.M.
receptors by molecular shape. J. Med. Chem., 1986, 29, 2149-2153. Structure-based discovery of inhibitors of thymidylate synthase.
[50] Kuntz, I.D.; Leach, A.R. Conformational analysis of flexible Science, 1993, 259, 1445-1450.
ligands in macromolecular receptor sites. J. Comput. Chem., 1992, [75] Michel, J.; Verdonk, M.L.; Essex, J.W. Protein-ligand binding
13, 730-748 affinity predictions by implicit solvent simulations: a tool for lead
[51] Ewing, T.J.; Makino, S.; Skillman, A.G.; Kuntz, I.D. DOCK 4.0: optimization? J. Med. Chem., 2006, 49, 7427-7439.
search strategies for automated molecular docking of flexible [76] Briggs, J.M.; Marrone, T.J.; McCammon, J.A. Computational
molecule databases. J. Comput. Aided Mol. Des., 2001, 15, 411- science new horizons and relevance to pharmaceutical design.
428. Trends Cardiovasc. Med., 1996, 6, 198-206.
[52] Welch, W.; Ruppert, J.; Jain, A.N. Hammerhead: fast, fully [77] Bohm, H.J. Prediction of binding constants of protein ligands: a
automated docking of flexible ligands to protein binding sites. fast method for the prioritization of hits obtained from de novo
Chem. Biol., 1996, 3, 449-462. design or 3D database search programs. J. Comput. Aided Mol.
[53] Schnecke, V.; Kuhn, L.A. Virtual Screening with solvation and Des., 1998, 12, 309-323.
ligand-induced complementarity. Persp. Drug Discov. Des., 2000, [78] Gehlhaar, D.K.; Verkhivker, G.M.; Rejto, P.A.; Sherman, C.J.;
20, 171-190. Fogel, D.B.; Fogel, L.J.; Freer, S.T. Molecular recognition of the
[54] Zsoldos, Z.; Reid, D.; Simon, A.; Sadjad, B.S.; Johnson, A.P. inhibitor AG-1343 by HIV-1 protease: conformationally flexible
eHiTS: an innovative approach to the docking and scoring function docking by evolutionary programming. Chem. Biol., 1995, 2, 317-
problems. Curr. Protein Pept. Sci., 2006, 7, 421-435. 324.
[55] Miranker, A.; Karplus, M. Functionality maps of binding sites: a [79] Verkhivker, G.M.; Bouzida, D.; Gehlhaar, D.K.; Rejto, P.A.;
multiple copy simultaneous search method. Proteins, 1991, 11, 29- Arthurs, S.; Colson, A.B.; Freer, S.T.; Larson, V.; Luty, B.A.;
34. Marrone, T.; Rose, P.W. Deciphering common failures in
[56] Eisen, M.B.; Wiley, D.C.; Karplus, M.; Hubbard, R.E. HOOK: a molecular docking of ligand-protein complexes. J. Comput. Aided
program for finding novel molecular architectures that satisfy the Mol. Des., 2000, 14, 731-751.
chemical and steric requirements of a macromolecule binding site. [80] Jain, A.N. Scoring non-covalent protein-ligand interactions: a
Proteins, 1994, 19, 199-221. continuous differentiable function tuned to compute binding
[57] Bohm, H.J. LUDI: rule-based automatic design of new substituents affinities. J. Comput. Aided Mol. Des., 1996, 10, 427-440.
for enzyme inhibitor leads. J. Comput. Aided Mol. Des., 1992, 6, [81] Head, R.D.; Smythe, M.L.; Oprea, T.I.; Waller, C.L.; Green, S.M.;
593-606. Marshall, G.R. VALIDATE: a new method for the receptor-based
[58] Goodsell, D.S.; Lauble, H.; Stout, C.D.; Olson, A.J. Automated prediction of binding affinities of novel ligands. J. Am. Chem. Soc.,
docking in crystallography: analysis of the substrates of aconitase. 1996, 118, 3959-3969.
Proteins, 1993, 17, 1-10. [82] Gehlhaar, D.K.; Moerder, K.E.; Zichi, D.; Sherman, C.J.; Ogden,
[59] Hart, T.N.; Read, R.J. A multiple-start Monte Carlo docking R.C.; Freer, S.T. De novo design of enzyme inhibitors by Monte
method. Proteins, 1992, 13, 206-222. Carlo ligand generation. J. Med. Chem., 1995, 38, 466-472.
[60] Goodsell, D.S.; Olson, A.J. Automated docking of substrates to [83] Eldridge, M.D.; Murray, C.W.; Auton, T.R.; Paolini, G.V.; Mee,
proteins by simulated annealing. Proteins, 1990, 8, 195-202. R.P. Empirical scoring functions: I. The development of a fast
[61] Abagyan, R.; Totrov, M.; Kuznetsov, D. ICM-A new method for empirical scoring function to estimate the binding affinity of
protein modeling and design: Applications to docking and structure ligands in receptor complexes. J. Comput. Aided Mol. Des., 1997,
11, 425-445.
156 Current Computer-Aided Drug Design, 2011, Vol. 7, No. 2 Meng et al.

[84] Muegge, I.; Martin, Y.C. A general and fast scoring function for [106] Teague, S.J. Implications of protein flexibility for drug discovery.
protein-ligand interactions: a simplified potential approach. J. Med. Nat. Rev. Drug Discov., 2003, 2, 527-541.
Chem., 1999, 42, 791-804. [107] Gschwend, D.A.; Good, A.C.; Kuntz, I.D. Molecular docking
[85] Mitchell, J.B.O.; Laskowski, R.A.; Alex, A.; Thornton, J.M. Bleep- towards drug discovery. J. Mol. Recognit., 1996, 9, 175-186.
potential of mean force describing protein-ligand interactions: I. [108] Totrov, M.; Abagyan, R. Protein-ligand docking as an energy
generating potential. J. Comput. Chem., 1999, 20, 1165 - 1176. optimization problem. In: Drug-receptor thermodynamics:
[86] Ishchenko, A.V.; Shakhnovich, E.I. SMall Molecule Growth 2001 Introduction and experimental applications; Raffa, R. B., Ed. John
(SMoG2001): an improved knowledge-based scoring function for Wiley & Sons: New York, 2001; pp. 603-624.
protein-ligand interactions. J. Med. Chem., 2002, 45, 2770-2780. [109] Leach, A.R. Ligand docking to proteins with discrete side-chain
[87] Feher, M.; Deretey, E.; Roy, S. BHB: a simple knowledge-based flexibility. J. Mol. Biol., 1994, 235, 345-356.
scoring function to improve the efficiency of database screening. J. [110] Desmet, J.; De Maeyer, M.; Hazes, B.; Lasters, I. The dead-end
Chem. Inf. Comput. Sci., 2003, 43, 1316-1327. elimination theorem and its use in protein sidechain positioning.
[88] Verkhivker, G.; Appelt, K.; Freer, S.T.; Villafranca, J.E. Empirical Nature, 1992, 356, 539-542
free energy calculations of ligand-protein crystallographic [111] Abagyan, R.; Totrov, M. Biased probability Monte Carlo
complexes. I. Knowledge-based ligand-protein interaction conformational searches and electrostatic calculations for peptides
potentials applied to the prediction of human immunodeficiency and proteins. J. Mol. Biol., 1994, 235, 983-1002.
virus 1 protease binding affinity. Protein Eng., 1995, 8, 677-691. [112] Morris, G.M.; Huey, R.; Lindstrom, W.; Sanner, M.F.; Belew,
[89] Wallqvist, A.; Jernigan, R.L.; Covell, D.G. A preference-based R.K.; Goodsell, D.S.; Olson, A.J. AutoDock4 and
free-energy parameterization of enzyme-inhibitor binding. AutoDockTools4: Automated docking with selective receptor
Applications to HIV-1-protease inhibitor design. Protein Sci., flexibility. J. Comput. Chem., 2009, 30, 2785-2791.
1995, 4, 1881-1903. [113] Knegtel, R. M.; Kuntz, I. D.; Oshiro, C. M., Molecular docking to
[90] Gohlke, H.; Hendlich, M.; Klebe, G. Knowledge-based scoring ensembles of protein structures. J. Mol. Biol., 1997, 266, 424-440.
function to predict protein-ligand interactions. J. Mol. Biol., 2000, [114] Carlson, H.A.; Masukawa, K.M.; Rubins, K.; Bushman, F.D.;
295, 337-356. Jorgensen, W.L.; Lins, R.D.; Briggs, J.M.; McCammon, J. A.,
[91] DeWitte, R.S.; Shakhnovich, E.I., SMoG: de novo design method Developing a dynamic pharmacophore model for HIV-1 integrase.
based on simple, fast, and accurate free energy estimates. 1. J. Med. Chem., 2000, 43, 2100-2114.
methodology and supporting evidence. J. Am. Chem. Soc., 1996, [115] Cavasotto, C.N.; Abagyan, R.A. Protein flexibility in ligand
118, 11733-11744. docking and virtual screening to protein kinases. J. Mol. Biol.,
[92] Charifson, P.S.; Corkery, J.J.; Murcko, M.A.; Walters, W.P. 2004, 337, 209-225.
Consensus scoring: a method for obtaining improved hit rates from [116] Derreumaux, P.; Schlick, T. The loop opening/closing motion of
docking databases of three-dimensional structures into proteins. J. the enzyme triosephosphate isomerase. Biophys. J., 1998, 74, 72-
Med. Chem., 1999, 42, 5100-5109. 81.
[93] Feher, M. Consensus scoring for protein-ligand interactions. Drug [117] Zeng, L.; Zhou, M.M. Bromodomain: an acetyl-lysine binding
Discov. Today, 2006, 11, 421-428. domain. FEBS Lett., 2002, 513, 124-128.
[94] Clark, R.D.; Strizhev, A.; Leonard, J.M.; Blake, J.F.; Matthew, J.B. [118] Venkitakrishnan, R.P.; Zaborowski, E.; McElheny, D.; Benkovic,
Consensus scoring for ligand/protein interactions. J. Mol. Graph. S.J.; Dyson, H.J.; Wright, P.E. Conformational changes in the
Model., 2002, 20, 281-295. active site loops of dihydrofolate reductase during the catalytic
[95] Srinivasan, J.; Cheatham, T.E.; Cieplak, P.; Kollman, P.A.; Case, cycle. Biochemistry, 2004, 43, 16046-16055.
D.A. Continuum solvent studies of the stability of DNA, RNA, and [119] Go, N.; Scheraga, H.A. Ring closure and local conformational
phosphoramidate−DNA helices. J. Am. Chem. Soc., 1998, 120, deformations of chain molecules. Macromolecules, 1970, 3, 178-
9401-9409. 187.
[96] Kollman, P.A.; Massova, I.; Reyes, C.; Kuhn, B.; Huo, S.; Chong, [120] Dodd, L.R.; Boone, T.D.; Theodorou, D.N. A concerted rotation
L.; Lee, M.; Lee, T.; Duan, Y.; Wang, W.; Donini, O.; Cieplak, P.; algorithm for atomistic Monte Carlo simulation of polymer melts
Srinivasan, J.; Case, D.A.; Cheatham, T.E. 3rd. Calculating and glasses. Mol. Phys., 1993, 78, 961-996.
structures and free energies of complex molecules: combining [121] Hoffmann, D.; Knapp, W. Polypeptide folding with off-lattice
molecular mechanics and continuum models. Acc. Chem. Res., Monte Carlo dynamics: the method. Eur. Biophys. J., 1996, 111,
2000, 33, 889-897. 387-404.
[97] Still, W.C.; Tempczyk, A.; Hawley, R.C.; Hendrickson, T. [122] Wu, M.G.; Deem, M.W. Analytical rebridging Monte Carlo:
Semianalytical treatment of solvation for molecular mechanics and application to cis/trans isomerization in proline-containing, cyclic
dynamics. J. Am. Chem. Soc., 1990, 112, 6127-6129. peptides. J. Chem. Phys., 1999, 14, 6625-6632.
[98] Guimaraes, C.R.; Mathiowetz, A.M. Addressing limitations with [123] Dinner, A.R. Local deformations of polymers with non-plannar
the MM-GB/SA scoring procedure using the WaterMap method rigid main chain internal coordinates. J. Comp. Chem., 2000, 21,
and free energy perturbation calculations. J. Chem. Inf. Model., 50, 1132-1144.
547-559. [124] Mezei, M. Efficient Monte Carlo sampling for long molecular
[99] Singh, N.; Warshel, A. Absolute binding free energy calculations: chains using local moves, tested on a solvated lipid bilayer. J.
on the accuracy of computational scoring of protein-ligand Chem. Phys., 2003, 118, 3874-3879.
interactions. Proteins, 2010, 78, 1705-1723. [125] Cui, M.; Mezei, M.; Osman, R. Prediction of protein loop
[100] Gabb, H.A.; Jackson, R.M.; Sternberg, M.J. Modelling protein structures using a local move Monte Carlo approach and a grid-
docking using shape complementarity, electrostatics and based force field. Protein Eng. Des. Sel., 2008, 21, 729-735.
biochemical information. J. Mol. Biol., 1997, 272, 106-120. [126] Kubinyi, H. Computer Applications in Pharmaceutical Research
[101] Bron, C.; Kerbosch, J. Algorithm 457: finding all cliques of an and Development John Wiley: New York, 2006.
undirected graph. Commun. ACM, 1973, 16, 575-576. [127] Kroemer, R.T. Structure-based drug design: docking and scoring.
[102] Meng, E.C.; Shoichet, B.K.; Kuntz, I.D. Automated docking with Curr. Protein Pept. Sci., 2007, 8, 312-328.
grid-based energy evaluation. J. Comput. Chem., 1992, 13, 505- [128] Venhorst, J.; ter Laak, A.M.; Commandeur, J.N.; Funae, Y.; Hiroi,
524. T.; Vermeulen, N.P. Homology modeling of rat and human
[103] Zou, X.Q.; Sun, Y.; Kuntz, I.D. Inclusion of solvation in ligand cytochrome P450 2D (CYP2D) isoforms and computational
binding free energy calculations using the generalized-born model. rationalization of experimental ligand-binding specificities. J. Med.
J. Am. Chem. Soc., 1999, 121, 8033-8043. Chem., 2003, 46, 74-86.
[104] Trott O, Olson AJ. AutoDock Vina: improving the speed and [129] Williams, P.A.; Cosme, J.; Ward, A.; Angove, H.C.; Matak
accuracy of docking with a new scoring function, efficient Vinkovic, D.; Jhoti, H. Crystal structure of human cytochrome
optimization, and multithreading. J. Comput. Chem., 2010, 31, 455- P450 2C9 with bound warfarin. Nature, 2003, 424, 464-468.
461. [130] Meng, X.Y.; Zheng, Q.C.; Zhang, H.X. A comparative analysis of
[105] Bohm, H.J. The development of a simple empirical scoring binding sites between mouse CYP2C38 and CYP2C39 based on
function to estimate the binding constant for a protein-ligand homology modeling, molecular dynamics simulation and docking
complex of known three-dimensional structure. J. Comput. Aided studies. Biochim. Biophys. Acta, 2009, 1794, 1066-1072.
Mol. Des., 1994, 8, 243-256.
Molecular Docking Current Computer-Aided Drug Design, 2011, Vol. 7, No. 2 157

[131] Boehm, H.J.; Boehringer, M.; Bur, D.; Gmuender, H.; Huber, W.; [134] Shoichet, B.K.; Leach, A.R.; Kuntz, I.D. Ligand solvation in
Klaus, W.; Kostrewa, D.; Kuehne, H.; Luebbers, T.; Meunier- molecular docking. Proteins, 1999, 34, 4-16.
Keller, N.; Mueller, F. Novel inhibitors of DNA gyrase: 3D [135] Lorber, D.M.; Shoichet, B.K. Flexible ligand docking using
structure based biased needle screening, hit validation by conformational ensembles. Protein Sci., 1998, 7, 938-950.
biophysical methods, and 3D guided optimization. A promising [136] Freymann, D.M.; Wenck, M.A.; Engel, J.C.; Feng, J.; Focia, P.J.;
alternative to random screening. J. Med. Chem., 2000, 43, 2664- Eakin, A.E.; Craig, S.P. Efficient identification of inhibitors
2674. targeting the closed active site conformation of the HPRT from
[132] Kirton, S.B.; Murray, C.W.; Verdonk, M.L.; Taylor, R.D. Trypanosoma cruzi. Chem. Biol., 2000, 7, 957-968.
Prediction of binding modes for ligands in the cytochromes P450 [137] Su, A.I.; Lorber, D.M.; Weston, G.S.; Baase, W.A.; Matthews,
and other heme-containing proteins. Proteins, 2005, 58, 836-844. B.W.; Shoichet, B.K. Docking molecules by families to increase
[133] Doman, T.N.; McGovern, S.L.; Witherbee, B.J.; Kasten, T.P.; the diversity of hits in database screens: computational strategy and
Kurumbail, R.; Stallings, W.C.; Connolly, D.T.; Shoichet, B.K., experimental evaluation. Proteins, 2001, 42, 279-293.
Molecular docking and high-throughput screening for novel [138] Gschwend, D.A.; Kuntz, I.D. Orientational sampling and rigid-
inhibitors of protein tyrosine phosphatase-1B. J. Med. Chem., 2002, body minimization in molecular docking revisited: on-the-fly
45, 2213-2221. optimization and degeneracy removal. J. Comput. Aided Mol. Des.,
1996, 10, 123-132.

Received: January 15, 2010 Revised: April 28, 2010 Accepted: September 29, 2010

You might also like