MACHINE LEARNING APPROACH
TO COMPUTATIONAL CHEMISTRY
SUBMITTED TO - SATYAM RAVI SIR
SUBMITTED BY
21BCG10123. Ambuj Singh
21BSA10094 Faysal Ezaz
21BAC10010. Ramji Chaurasiya
21BSA10127. Smarika Malviya
21BSA10123. Nikita Raut
21BAI10484 Ishali Dubey
21BSA10133. Vansh Thakur
21BHI10028 Arnab Choudhry
21BSA10109 Ayush Sharma
21BCG10128 P S Samendh Krishna
21BCG10143 Tushar Sharma
21MIP10014 Priyangshu Das
MACHINE LEARNING
Machine learning (ML) is a type of arti icial intelligence (AI) that allows
software applications to become more accurate at predicting outcomes
without being explicitly programmed to do so. Machine learning algorithms
use historical data as input to predict new output values.
f
MACHINE LEARNING IN
CHEMISTRY
Molecular simulations provide deep insight into chemical processes beyond what can be
directly measured experimentally, holding major promise for accelerating the discovery of
molecules and materials.
Theoretical and computational modelling is ubiquitous in materials research.
Modelling can signi icantly help to bridge the results of fundamental materials
research to actual materials production by signi icantly reducing timescales. The
computational chemistry approaches developed over the years have been an
invaluable tool to provide deep insight into chemical processes beyond what can be
f
directly measured experimentally.
f
Computational studies of chemical processes taking place over extended
size and time scales must balance computational cost and accuracy:
electronic structure methods are very accurate but computationally
expensive, while atomistic models such as force ields—although
computationally affordable—lack transferability to new systems.
f
Drug design
What is drug design?
Drug design is the inventive process of inding new medications based on the
knowledge of a biological target.
Multi Disciplinary process that includes
1. Identi ication of a Drug target
2. Bioassays
3. Structural Biology
4. Synthetic Medicinal Chemistry
5. Evaluation of the drug behaviour in the body
f
f
Computer Aided Drug Design (CADD)
Drug design which relies on computer modelling techniques and machine learning
techniques is referred to as computer aided drug design
Computer aided drug design uses computational chemistry + machine learning to
discover, enhance or study drugs and related biologically active molecules
Ligand Based Drug Design
Pharmacophore model that de ines the minimum necessary structural characteristics a molecule must possess
in order to bind to the target.
model of the biological target may be built based on the knowledge of what binds to it, and this model in turn
may be used to design new molecular entities that interact with the target.
A model of the biological target may be built based on the knowledge of what binds to it, and this model in turn
may be used to design new molecular entities that interact with the target.
f
Receptor-based drug design :
• Structure-based drug design (or direct drug design) relies on knowledge
of the three dimensional structure of the biological target obtained
through methods such as x-ray crystallography or NMR spectroscopy.
If an experimental structure of a target is not available, it may be possible
to create a homology model of the target based on the experimental
structure of a related protein.
Contributions of machine
learning in drug designing
The ML model was used to search through thousands of approved ligands by the
Food and Drug Administration (FDA) and a million biomolecules in the BindingDB
database. From these, insights were obtained for more than 19,000 molecules
satisfying the Vina score (i.e. an important physicochemical measure of the
therapeutic process of a molecule that is used to rank molecular conformations and
predict free energy of binding)
Machine learning for chemical discovery
Computational design and discovery of molecules and materials relies on
the exploration of increasingly growing chemical space. Obviously,
chemical discovery concerns not only with inding “this special molecule”,
but also predicting reaction pathways and interactions between molecules,
optimizing catalytic conditions, eliminating undesired side effects, among
many other important degrees of freedom.
f
Chemical discovery and ML are bound to evolve together, but achieving true
synergy between them requires solving many outstanding challenges. The potential
of using ML for increasing the accuracy and ef iciency of molecular simulations has
been established beyond any doubt Data-driven high-throughput materials
discovery has also been established as a ield of its own.
f
f
Physically inspired ML algorithms can identify new drug candidates, ind new phases
in amorphous materials, carry out molecular dynamics with essentially exact
quantum forces, and offer unprecedented statistical insights into chemical
environments. Up to now, most of these applications were done under idealized
conditions.
f
From molecular big data to
chemical discovery
The quality and reliability of ML models in any scienti ic domain depends on the
increasing availability of [Link] development of physics-inspired ML models and
sophisticated atomistic descriptors have been crucial for increasing the predictive
power of ML models by at least two orders of magnitude in the past years an
incredible scienti ic progress.
f
f
Some successful ML-driven discoveries have been made in the search for organic
light-emitting diodes, redox- low batteries, and antibiotics, among many other
examples. The most remarkable aspect of ML for chemical discovery is that the
corresponding statistical view on chemical space often enables asking new
questions and obtaining novel insights.
f
Future of ML for chemical
discovery
Current successful applications of ML for chemical discovery have only scratched
the surface of possibilities. There are many conceptual, theoretical, and practical
challenges waiting to be solved to enable the “chemical discovery revolution”. Here I
discuss the challenges that I consider to be the most pressing and interesting at this
moment.A universal ML approach should have the capacity to accurately predict
both energetic and electronic properties of molecules. In addition, such an
approach should uniformly describe compositional (chemical arrangement of atoms
in a molecule) and con igurational (physical arrangement of atoms in space) degrees
of freedom on equal footing.
f
Validation of ML predictions ultimately requires comparison to experimental
observables, such as reaction rates, spectroscopic observations, solvation energies,
melting temperatures, among other relevant quantities. Calculating these
observables demands a tight integration of QM, statistical simulations, and fast ML
predictions, all integrated in a comprehensive molecular simulations framework.
Solving many of the challenges posed above will require coming up with creative
interdisciplinary approaches combining quantum and statistical mechanics,
chemical knowledge, and sophisticated ML tools, irmly based on growing datasets
that cover increasingly broader domains of the vast chemical space.
f
Thermodynamics And Machine
learning
Many Comp Chemistry efforts focus on predicting thermodynamic properties at
inite temperatures, such as heat capacity, density, and chemical potential. Although
many physical properties are already accessible from MD simulations, doing
estimations of free energies that establish the relative stability of different states
using electronic structure methods re-mains dif icult.
f
f
The con igurational part of the Gibbs free energy of a bulk system that has N
distinguishable particles with atomic coordinates r = {r1...N }, and the associated
potential energy. In order to rigorously determine G, one must exhaustively sample
the con iguration space that has relatively high weight exp. This normally requires
thermodynamics integration or enhanced sampling methods (e.g. umbrella
sampling, metadynamics, transition
path sampling, that require simulation times and scales far beyond what is
accessible with MD simulations based on KS-DFT or correlated wavefunction
methods.
f
f
However, MLPs have unleashed both limits on the time scale and system size. An
early example, used an MLP with umbrella sampling and the free energy
perturbation method to reveal the in luence of van der Waals corrections on the
thermodynamic properties of liquid water
f
the combination of an MLP trained from hybrid DFT data and free energy methods
reproduced several thermodynamic properties of water from quantum mechanics,
including the density of ice and water, the difference in melting temperature for
normal and heavy water, and the stability of different forms of ice. employed the
DeepMD approach to study the relatively long time-scale nucleation of gallium.
MLPs for high-pressure hydrogen provided evidence on how hydrogen gradually
turns into a metal in giant planets. In all these examples, high accuracy and long
timescales were required to model the speci ic phenomena and reveal physical
insights, and it is precisely the combination of CompChem+ML that enables both.
f
Retrosynthetic technologies
A grand challenge in chemistry is to understand synthetic pathways to desired
molecules.
Retrosynthesis involves the design of chemical steps to produce molecules and
materials that would be crucial to drug discovery, medicinal chemistry, and materials
science.
First, simple combinatorics make the space of possible reactions greater than the
space of possible molecules.
Second, reactants seldom contain only one reactive functional group, and thus
require predictions of multiple functional groups.
Third, one failed step in the route can invalidate the entire synthesis because organic
synthesis is a multistep process.
Given these challenges, ML is becoming more established in determining reaction
rules from computational chemistry data.
These programs are typically based on one of three possible algorithms:
1. Algorithms that use reaction rules (manually encoded or automatically derived
from
databases).
2. Algorithms that use principles of physical chemistry based on ab initio
calculations to
predict energy barriers.
3. Algorithms based on ML techniques
ML approaches are used to try to overcome the generalization issues of rule-based
algorithms (that normally suffering from incompleteness, infeasible suggestions, and
human bias) while also avoiding the high cost of Computational Chemistry
calculations.
It is now possible to obtain purely data-driven approaches for synthesis planning,
which are promoting a rapid advancement in the ield. . For example, Coley and co-
workers designed a data-driven metric, SC score, for describing a real synthesis
modeled after the idea that products are, on average, more synthetically complex
than each of their reactants.
f
Catalysis
Catalysis research requires multiscale approaches to determine chemical
compounds that can in luence barrier heights of reaction mechanisms to impact
product yields and selectivities without otherwise being generated or consumed by
the reaction. Traditional catalysis is normally discussed in textbooks in terms of
homogeneous (i.e. within a solution phase), heterogeneous (occurring at a solid/
liquid interface), and biological (occurring within enzymes and ribo enzymes), but it
is best not to use these terms too strictly because actual reaction mechanisms can
be quite complex and overall processes may sometimes exhibit characteristics (by
design) of two or more of these classical processes.
f
• Catalysis makes up roughly 35% of the world’s gross domestic product, 655 and it i
important to guide toward the end goal of achieving greater sustainability with
catalytic processes.
These reasons help make catalysis a fertile training ground for applying and
developing theoretical models that can be used along with Computational
Chemistry or computational chemistry + ML.
The research ield is also burgeoning with many reports and review articles that
discuss perspectives and progress using ML methods for catalysis science; here, w
will mention notable examples that present a broad range of ways that Comp
Chemistry +ML can be used for insights.
f
For example, Comp Chemistry + ML methods are enabling more data generation by
allowing costly QC calculations to be run more ef iciently, and more information
means more comprehensive predictions of chemical and materials phase diagrams
for catalysis as well as stability and reactivity descriptors are identi ied.
ML approaches can be used in catalysis studies to deduce insights into interaction
trends between single metal atoms and oxide supports, to identify the signi icance
of features (e.g. adsorbate type or coverage) where Comp Chemistry theories break
down , or they can be used to identify trends that result in optimal catalysis across
multiple objectives such as activity and cost.
f
f
f
Computational Chemistry And
Machine Learning
ML is also opening opportunities for Comp Chemistry+ ML studies on highly detailed
and complex networks of reactions. Such models in principle can then signi icantly
extend the range of utility of micro kinetics modeling for predictions of products
from catalysis.
ML also enables studies of complicated reaction networks that can allow predictions
of region selective products based on Comp Chemistry data, asymmetric catalysis
important for natural product synthesis , and biochemical reactions.
f
AI for chemistry
Arti icial Intelligence (AI) is being used increasingly by chemists to perform various
tasks. Originally, research in AI applied to chemistry has been fueled by the need to
accelerate drug discovery and reduce its huge costs and the time to market for new
drugs.
f
Molecule property prediction:
When scientists design new molecules for a certain application, they must
synthesize them to check experimentally that they possess the right properties. If
they do not, the scientists design new molecules (that can be analogs of the
previously synthesized molecules, for example), and iterate until they obtain
molecules that satisfy their requirements (properties, performance, price, toxicity,
environmental impact, etc.).
This iterative process takes a lot of time and money. Being able to accurately predict
the properties of hypothetical molecules would allow researchers to synthesize only
the most promising ones and to avoid synthesizing and testing many molecules that
do not possess the desired properties. Methods of prediction of molecule properties
have been used for a long time, often under the name of Quantitative Structure-
Activity Relationships (QSAR).
Molecule design
Designing new molecules is one of the highest value-added tasks that chemists
perform. They usually use their chemistry knowledge, their domain knowledge, and
their creativity to propose new molecular structures that are then tested either
virtually (in silico) or experimentally in the relevant applications.
There are two main limitations to this method. The irst limitation is that human
creativity is inherently biased: a chemist may prefer certain functional groups or
molecular patterns (consciously or not), and might exclude molecules that he/she
inds strange or does not like.
The second limitation is that when large numbers of new molecules (hundreds of
thousands or even millions) are needed for virtual screening, the human mind is not
capable of designing such large sets of molecules in a reasonable time.
f
f
Chemical reaction optimization
Once you have determined what molecule you want to synthesize and which
reaction steps you want to use to synthesize it, the next step is to go in the
laboratory and experimentally search for synthesis conditions. For example, you
want to know the temperature, concentrations, quantity of catalyst and reaction
time that will lead to the desired molecule in the highest possible yield. This search
for optimal reaction conditions can be exceptionally long and require many
synthesis attempts, especially if there are many different parameters to control and
that you change only one of them at each experiment ("one variable at a time"
method).
THANK YOU