Computer-aided drug design
Known ligand(s) No known ligand
Known protein
Structure-based drug
design (SBDD) De novo design
structure
Protein-ligand docking
Ligand-based drug design
protein structure
(LBDD)
1 or more ligands CADD of no use
• Similarity searching Need experimental
Unknown
Several ligands data of some sort
• Pharmacophore searching
Many ligands (20+)
• Quantitative Structure-Activity
Relationships (QSAR)
Docking
• Docking refers to a computational scheme that tries
to find the best binding orientation between two
biomolecules where the starting point is the atomic
coordinates of the two molecules
Structure-based drug design (SBDD)
PDB: Protein Data Bank [Link]
Macromolecule
With Co-crystal inhibitor/activator Without inhibitor
Known Binding Site Unknown Binding Site
Active site directed Docking Blind Docking
Stages of Docking
Pose generation
–Place the ligand in the binding site
–Generally well solved
Algorithms
A process or set of rules to be followed in calculations or other problem-solving operations, especially by a computer.
Pose selection
–Determine geometry of the ligand
Hard to find Scoring function
Geometry = location, orientation and
conformation
Approaches for docking
O H
Representing — Searching — Evaluating
• Surface representations • Rule-based • Force fields
• Physicochemical descriptors • Energy-driven • Regression functions
• Grid-based approaches • Knowledge-based potentials
Representing: Molecular representations for docking
A) Protein
Restricting the search space to the binding pocket
Geometric surface descriptors
DOCK: sphere representation
Physicochemical descriptors O H
LUDI, FlexX: Interaction points
and vectors
Grid representations
Pg ,t Wt ,T ( p ) (r )
Autodock pP
Interaction potentials of probe g: grid point
t: ligand atom type (probe)
T(p): atom type of protein atom p r: distance
atoms are mapped to grid
points
Different types of docking strategies
Blind docking: No idea about Binding site.
Knowledge based docking: Known Binding site.
Rigid docking
Flexible docking
Bound docking: reproduce a known complex where the starting point
is atomic structures from a co-crystal.
Unbound docking: docking of newly designed molecules
Docking Methods
Rigid body docking
Flexible-ligand docking
Flexible docking
Fast, Simple
•Initially –Receptor (protein) and ligand rigid
•Most current approaches –Receptor rigid, ligand flexible
•Advanced approaches–Receptor (to a degree) and ligand flexible
Slow, Complex
Docking Methods
Rigid body docking – protein remains fixed, small molecule has 6
degrees of freedom (DOF) – 3 translational and 3 rotational
Docking Methods
Flexible-ligand docking – protein remains fixed, small molecule
has standard 6 DOF plus internal DOF – can rotate about bonds
– More time consuming, but necessary for complex ligands if
binding conformation is unknown
Flexible docking – as above, and in addition protein atoms in
neighbourhood of binding site can move
– Largest conformational space to search
– Often done by using multiple static protein conformers, and
treating each by flexible ligand docking
– Often important when docking to apo-protein e.g. allosteric
effects
Basic Principles
• The association of molecules is based on interactions
– Hydrogen bonds, salt bridges, hydrophobic contacts
– Electrostatics
– Very strong repulsive interactions on short distances (van der Waals)
•The associative interactions are weak and short-range
–Tight binding implies surface complementarity
•Most molecules are flexible
•Macro molecules are restricted in conformational space in a complicated way
Terminology: conformation, configuration, pose
Conformation: the relative positions of atoms in the 3D
structure of a molecule, independent of the coordinate system
2 different conformations of a ligand
Terminology: conformation, configuration, pose
Configuration/placement: the positions of atoms
of a molecule after undergoing a rigid transformation
(rotation and translation) in a coordinate system
2 different configurations (of same conformation)
Terminology: conformation, configuration, pose
Pose: a configuration of a conformation of a molecule
in a coordinate system
2 different poses of a ligand
Representations of Molecular Conformation
Cartesian Coordinates
e.g. PDB, Mol2
Distance Matrix
Internal Coordinates
– Bond length, bond angle, torsion angle
– E.g. Z-Matrix
Searching methods
Searching conformational space during docking:
1. Monte carlo method:
rigid-body translation or rotation.
energy- based selection criterion
Random moves and accept/reject based on Boltzmann probability.
2. Simulated annealing: Based on temperature effects
Starts with high temp-> global search
Low temp->local search
P=e(-ΔE/KT) { ΔE=E1-E0}
3. Genetic algorithm:
Darwin theory : Genotype->phenotype
Lamarckian theory: Phenotype->Genotype
E.g.: GOLD,Autodock,GLIDE
4. Matching algorithms (MA) :
Based on molecular shape map a ligand into an active site of a protein in terms of
shape features and chemical information(Pharmacophore).
Distance matrix between the pharmacophore and the corresponding ligand atoms.
Chemical properties, like hydrogen-bond donors and acceptors, can be taken into
account during the match.
Advantage : speed;
Used for the enrichment of active compounds from large libraries.
Simulated Annealing
• Simulates how the real systems reach their minimum energy
forms.
• In simulated annealing methods, a molecule is being subject to
a very high temperature during an molecular dynamics or an
Monte Carlo search
• It is then cooled down slowly, allowing it to be trapped in a local
minimum
• The process can be repeated if desired
Towards a minimum energy conformation
•The temperature of the system is decreased over time, until
a stable docked position is obtained.
•This allows the ligand to end up lower energy structure in
the bound state and drives the whole system towards a
minimum energy.
Genetic Algorithms
Scoring Functions
Scoring function
Forcefield-based
• Based on terms from molecular mechanics force fields
• usually quantify the sum of two energies, the receptor–ligand
interaction energy and internal ligand energy (such as steric
strain induced by binding).
• GoldScore, DOCK, AutoDock
Empirical
• Parameterised against experimental binding affinities.
• fit to reproduce experimental data, such as binding
energies and/or conformations
• ChemScore, PLP(piecewise linear potential), Glide SP/XP
Knowledge-based potentials
• Fit to reproduce experimental structures rather than binding
energies
• Based on statistical analysis of observed pairwise distributions
• PMF (potentials of mean force), DrugScore
Force Field: First used by spectroscopists to mean a set of equations designed to
reproduce or predict vibrational spectra.
A mathematical expression relating the potential energy of a molecule on the 3-
dimensional disposition of nuclear coordinates of the constituent atoms.
FF is a set of functions and constants used to find the potential Energy of the
molecule
The data determined experimentally for small molecules can be extrapolated to
larger molecules
The potential energy of a molecule can be written
E
bonds
stretch E
angles
bend Etorsions Enonbond
dihedrals pairs
Scoring Functions
Molecular Mechanics Scoring Functions
Chemistry at HARvard Macromolecular Mechanics
cross-term accounting for angle bending using 1,3 nonbonded interactions
Molecular Mechanics Scoring Functions
Assisted Model Building with Energy Refinement
Scoring Functions
Empirical Scoring Functions
entropy penalties on binding from a
weighted sum of the number of rotatable bonds in ligands.
ChemScore implements ligand rotational entropy
in a more complicated form that describes the molecular
environment surrounding each rotatable bond
Empirical Scoring Functions
Empirical Scoring Functions
Scoring Functions
Knowledge-Based Scoring Functions
Scoring Functions
Computing Scoring Functions
Computing Scoring Functions
Docking Methodology
Ligand Protein
Sketch or download Download from PDB or generate
convert to 3D homology model
Add Hydrogens
Add Hydrogens
Energy minimization Remove waters with in 5A0 distance
and non-amino acid elements , retain
the inhibitor
Energy Minimization
Define grid co-ordinates around binding site
Docking(Select the method of docking and algorithm)
Selection of compound based on binding affinity, interacting amino acids.
15/11/2014 49
Rigid Docking
Both molecules rigid
Orientation change: translation, rotation
No change in conformation
Shape complementary based on protein.
LOCK and KEY model
The DOCK algorithm – Rigid docking
• The DOCK algorithm developed
by Kuntz and co-workers is
generally considered one of the
major advances in protein–ligand
docking [Kuntz et al., JMB, 1982,
161, 269]
• The earliest version of the DOCK
algorithm only considered rigid
body docking and was designed to
identify molecules with a high
degree of shape complementarity
to the protein binding site.
• The first stage of the DOCK
method involves the construction
of a “negative image” of the
binding site consisting of a series
of overlapping spheres of varying
radii, derived from the molecular
surface of the protein
AR Leach, VJ Gillet, An Introduction to Cheminformatics
• Ligand atoms are then matched to
the sphere centres so that the
distances between the atoms
equal the distances between the
corresponding sphere centres,
within some tolerance.
• The ligand conformation is then
oriented into the binding site. After
checking to ensure that there are
no unacceptable steric
interactions, it is then scored.
• New orientations are produced by
generating new sets of matching
ligand atoms and sphere centres.
The procedure continues until all
possible matches have been
considered.
Flex X docking
Induced fit docking
Docking Algorithms
Stochastic Search:
– Genetic Algorithm, Monte Carlo simulated
annealing
– AutoDock, MCDock, ICM, GOLD, Glide
Incremental Construction:
– Rigid fragments with rotatable bonds
– Incremental : preferred torsion angles
– DOCK, FlexX, SLIDE, Surflex
Multiconformer:
– Generate a set of low-energy conformers
– Rigid docking
– FLOG, FRED, Yucca
Algorithms used while docking
• Fast shape matching (e.g., DOCK and Eudock)
• Incremental construction (e.g., FlexX, Hammerhead, SLIDE)
• Tabu search (e.g., PRO_LEADS and SFDock)
• Genetic algorithms (e.g., GOLD, AutoDock, and Gambler)
• Monte Carlo simulations (e.g., MCDock and QXP)
Searching conformational space before docking: (induced fit docking)
•Low energy conformation->rigid placement in binding site
•E.g.: SLIDE (Screening for ligands by induced fit docking)
•FRED (Fast rigid Exhaustive docking)
Incremental docking (Fragment based approach)
•E.g.: Flex X,HOOK
Multiple Copy Simultaneous Search (MCSS):
1,000 to 5,000 copies of a functional group, which are randomly placed in the binding
site of interest and subjected to simultaneous energy minimization and/or quenched
molecular dynamics in the forcefield of the protein
ALGORITHM and SCORING FUNCTION USED IN DOCKING
– AutoDock
• Uses a genetic algorithm
• Energy-based scoring
– GOLD
• Genetic algorithm
• Modified energy-based scoring
– DOCK
• Geometry-based
• Choice of scoring function
– FlexX
• Geometry based
• Uses Bohn scoring function
– ICM
• Monte Carlo minimization
• Scored based on ligand-ligand and protein-ligand interaction
– 3D-Dock [Link]
• which uses an unusual “Fourier correlation” method and is
aimed at protein-protein interactions
Tabu Search
• Combination of a minimization procedure with restrictions on the
search path
• A record of all the docked conformations will be stored
• The results of the subsequent steps would be either taken or
discarded based on the comparison
• As you move in the search space, keep a list of solutions and
their locations
• This algorithm doesn’t let your search, which has been
considered before