Unit VNotes
Unit VNotes
UNIT-V
COMPUTATIONAL CHEMISTRY
Abstract
Molecular modelling has emerged as a powerful tool in modern chemistry, offering cost-effective
and efficient alternatives to experimental approaches for understanding molecular structure and
behavior. This chapter introduces the scope of computational modelling, emphasizing its role in
predicting molecular interactions and properties. Key concepts of molecular topology are
explored, highlighting methods of structural representation through topological matrices such as
vertex adjacency, edge adjacency, and distance matrices. The discussion extends to topological
indices, including the Zagreb index, Wiener number, and Platt number, which serve as
quantitative measures of molecular structure and connectivity. By linking these indices with
quantitative structure–property and structure–activity relationships (QSPC/QSAR), the chapter
demonstrates how molecular modelling can aid in rational drug design, materials science, and
chemical property prediction. Overall, the integration of computational efficiency, structural
representation, and predictive analytics underscores the transformative potential of molecular
modelling in research and innovation.
Blow-up Syllabus
Sl No Topic to be taught Duration
1 Introduction to molecular modelling: Scope, cost and efficiency of 1 hour
computational modelling
2 Molecular interactions: Electrostatic interactions and Hydrogen bonding 1 hour
3 Molecular topology, topological matrix representation (vertex adjacency, 1 hour
edge adjacency and distance)
4 Hands on activities on topological matrices 1 hour
5 Topological indices (Zagreb, Wiener, Platt number) 1 hour
6 Hands on activities on topological indices 1 hour
7 Prediction of molecular properties through QSPC/QSAR. 1 hour
8 Hands on activities on QSPC/QSAR 1 hour
9 Case studies on QSAR/QSPC 1 hour
After going through this chapter, student will
(i) be able to understand the scope, advantages, and limitations of molecular
modelling in terms of cost, efficiency, and predictive capabilities compared to
experimental methods.
(ii) be able to Apply molecular topology concepts and construct topological
matrices (vertex adjacency, edge adjacency, and distance matrices) to represent
molecular structures.
(iii) be able to evaluate and utilize topological indices such as the Zagreb index,
Wiener number, and Platt number to predict molecular properties and relate
them to QSPC/QSAR approaches.
5.1 Introduction
With technological advancements, research has become faster through computational tools and
software, often used at the initial stage before experiments. Interdisciplinary research benefits
from combining computing skills with domain knowledge, though building and integrating
these skills remains a challenge. Earlier, computational chemistry was limited to experts due
to complex tools, but modern software is more user-friendly, supported by extensive literature
ranging from basic texts to advanced research papers. The main challenge now lies with
practicing chemists who wish to apply these techniques to real-world problems without delving
deeply into theory. These are experienced researchers who know chemistry and now have
computational tools available. These are people who want to use computational chemistry to
address real-world research problems and are bound to run into significant difficulties. This
unit is chosen to cover a large number of topics, with an emphasis on when and how to apply
computational techniques rather than focusing on theory. It gives a clear description with just
the amount of technical depth typically necessary to be able to apply the techniques to
computational problems. There are many good books describing the fundamental theory on
which computational chemistry is built. The description of that theory as given here is very
minimal. We have chosen to include just enough theory to explain the terminology used in
computing.
Chemistry impacts society through materials such as foods, fuels, drugs, and electronics, where
theory complements rather than replaces experiment. Their synergy can be grouped into three
categories: (1) post-experiment analysis, where theory resolves ambiguities in interpreting
results, e.g., predicting spectra to identify products. (2) Simultaneous use, where theoretical
predictions guide experiments, optimize design, and help detect minor components. (3)
Predictive applications, where theory estimates properties that are difficult, costly, or
dangerous to measure, such as atmospheric reaction rates or screening toxic/explosive
compounds. In this way, computational methods enhance efficiency, safety, and accuracy in
chemical research.
The cost and efficiency of computational modelling make it a highly valuable tool in modern
science and engineering. Compared to traditional experimental methods, computational
modelling significantly reduces the financial burden associated with chemicals, equipment,
laboratory facilities, and manpower. For example, expensive reagents, rare catalysts, or
hazardous substances can be studied virtually, avoiding waste and safety risks. Large-scale
screening of thousands of compounds in drug discovery or materials research can be carried
out on computers at a fraction of the cost of laboratory trials. In terms of efficiency,
computational modelling accelerates research by enabling rapid simulations and predictions
that would take much longer through experiments. It allows parallel testing of multiple
hypotheses, optimization of reaction conditions, and prediction of molecular properties before
synthesis. Modern high-performance computing (HPC) and cloud-based platforms further
enhance efficiency by handling complex simulations in shorter times. Moreover, computational
models can provide detailed atomic and molecular-level insights that are often inaccessible
experimentally, making research more precise and targeted. In summary, computational
modelling offers high cost-effectiveness and efficiency by saving time, reducing expenses,
minimizing risks, and enabling high-throughput virtual experimentation, while also
complementing laboratory studies to achieve faster and more reliable outcomes.
Molecular interactions are attractive or repulsive forces between molecules and between non-
bonded atoms. Molecular interactions are important in all aspects of chemistry, biochemistry
and biophysics, including protein folding, drug design, pathogen detection, material science,
sensors, gecko feet, nanotechnology, separations, and origins of life. Molecular interactions
are also known as noncovalent interactions, intermolecular interactions, non-bonding
interactions, noncovalent forces and intermolecular forces. All of five of these phrases mean
the same thing.
Bonding Interactions. Bonds hold atoms together within molecules. A molecule is a group of
atoms that associates strongly enough that it does not dissociate or lose structure when it
interacts with its environment. At room temperature two nitrogen atoms can be bonded (N 2).
Bonds break and form during chemical reactions. In the chemical reaction called fire, bonds
of cellulose break while bonds of carbon dioxide and water form. Bond enthalpies are on the
order of 100 kcal/mole (400 kjoule/mole), which is much greater than RT at room temperature;
bonds do not break at room temperature.
Boiling Points. When a molecule transitions from the liquid to the gas phase (as during
boiling), ideally all molecular interactions are disrupted. Ideal gases are the ONLY systems
where there are no molecular interactions. Differences in boiling temperatures give good
qualitative indications of strengths of molecular interactions in the liquid phase. High boiling
liquids have strong molecular interactions. The boiling point of H2O is hundreds of degrees
greater than the boiling point of N2 because of stronger molecular interactions in H2O (liq)
than in N2(liq). The forces between molecules in H2O (liq) are greater than those in N2(liq).
Atoms take space. Force two atoms together and they will push back. When two atoms are
close together, the occupied orbitals on the atom surfaces overlap, causing electrostatic
repulsion between surface electrons. This repulsive force between atoms acts over a very short
range, but is very large when distances are short. The repulsive energy goes up as (di / R)12,
where R is the distance between the atoms and di is the distance threshold below which the
energy becomes repulsive. di depends on the types of atoms. The large exponent means that
when R < di then small decreases in R cause large increases in repulsion. Short range repulsion
only matters when atoms are in very close proximity (R < di), but at close range it dominates
other interactions. Because this repulsion rises so sharply as distance decreases it is often useful
to pretend that atoms are hard spheres, like very small pool balls, with hard surfaces (called
van der Waals surfaces) and well- defined radii (called van der Waals radii). As two atoms
approach each other their van der Waals surfaces make contact when the distance between
them equals the sum of their van der Waals radii. At this distance the repulsive energy
skyrockets. The smallest distance between two non-bonded atoms is the sum of the van der
Waals radii of the two atoms. A sulfur atom and a carbon atom can come no closer together
than:
Of course, we are assuming here that bonds do not form. When two atoms form a bond, they
come very close together and their der Waals radii and surfaces are violated.
Figure 5.1 shows how short-range repulsion sets the distance of 3.4 Å between sheets in
graphite. If two non-bonded atoms are separated by a distance of less than the sum of their
VDW radii, short range repulsion forces them apart.
Short range repulsion is important to you. It prevents your hands from passing through each
other when you clap, and prevents atoms from collapsing into tightly packed states of
enormous density of 1014 g/ml, which is the density of condensed atomic nuclei. Here in earth,
with our modest gravity, the van der Waals radius of carbon (rC) is evident from the spacing
between the layers in graphite. The distance between atoms in different layers of graphite is
never less than twice the van der Waals radius of carbon (2 x rC = 2 x 1.7 = 3.4 Å). The atoms
within a graphite layer are covalently linked (bonded), which causes interpenetration of van
der Waals surfaces. Carbon atoms within a layer are separated by 1.42 Å, which is much less
than twice the van der Waals radius of carbon. As explained in other sections of this document
vdw surfaces are also violated when molecules form hydrogen bonds.
Electrostatic interactions:
Electrostatic interactions are between and among cations and anions, species with charge of
...-2, -1, +1, +2... Electrostatic interactions can be either attractive or repulsive, depending on
the signs of the charges. Like charges repel. Unlike charges attract. Favorable electrostatic
interactions cause the vapor pressure of sodium chloride and other salts to be very low. If you
leave crystals of table salt (NaCl; Na+=cation, Cl-=anion) on a hot pan, how long does it take
before they vaporize and sublime away? A very very long time; electrostatic interactions are
very very strong. The electrostatic interactions within a sodium chloride crystal are called ionic
bonds. But when a single cation and a single anion are close together, within a protein, or
within a folded RNA, those interactions are considered to be non-covalent electrostatic
interactions. Non-covalent electrostatic interactions can be strong, and act at long range.
Electrostatic forces fall off gradually with distance (1/r2, where r is the distance between the
ions).
Figure 5.2 shows electrostatic interactions in a cross section of a NaCl crystal. Each sodium
cation experiences strong electrostatic interactions with adjacent chloride anions.
Figure 5.3 shows electrostatic interactions. In RNA (for example in the ribosome), anionic
phosphate oxygens (charge = -1) engage in attractive electrostatic interactions with a
magnesium cation (charge = +2). Two phosphate groups can 'clamp' onto the Mg2+ ion. The
O to Mg2+ distance is 2.1 Å. The dashed lines represent favorable electrostatic interactions.
Electrostatic interactions are the primary stabilizing interaction between phosphate oxygens
of RNA (charge = -1) and magnesium ions (charge = +2), as shown in the figure below. There
are many magnesium ions associated with RNA and DNA in vivo. Electrostatic interactions
are highly attenuated (dampened) by water. In protein folding, RNA folding and DNA
annealing, electrostatic interactions are dependent on salt concentration and pH.
Favorable electrostatic interactions between paired anionic and cationic amino acid sidechains
are reasonably frequent in proteins. Ion Pairs, sometimes called Salt Bridges, are formed when
the charged group of a cationic amino acid (like lysine or arginine) is around 3.0 to 5.0 Å
from the charged group of an anionic amino acid (like aspartate or glutamate). The charged
groups in an ion pair are generally linked by hydrogen bonds, in addition to electrostatic
interactions.
Force = k q1 q2 / ε r2
where k = 9.0 x 109 nt-meter2 / coul2, q = -1.6 x 10-19 coulombs for an electron., r = distance
between the point charges (meters) and ε = the dielectric constant of the medium (unitless). It
reflects the tendency of the medium to shield charged species from each other. ε is 1 in a
vacuum, around 4 in the interior of a protein and 80 in water. Water is very efficient at shielding
charges, reducing electrostatic forces between ions. The problem of calculating electrostatic
effects in biological systems is complex in part because of non-uniformity of the dielectric
environment. The dielectric micro-environments are complex and variable, with less shielding
of charges in regions of hydrocarbon sidechains and greater shielding in regions of polar
sidechains. The electrostatic energy is given by:
𝑘𝑎𝑞1 𝑞2
∆𝐸 = , where a =Avogadro’s number.
∈𝑟
One can crudely estimate the energetics of a charge-charge interaction in a protein. The energy
of an amine (charge +1) and a carboxylic acid (charge -1) separated by 4 Å in the interior of
protein is given by:
= 87 kjoule/mole = 21 kcal/mole
This rough approximation is around 10-fold greater than the values determined experimentally.
An ion pair contributes favourable ΔG of 1 to 4 kcal/mole (4.1 to 16.4 kjoule/mole) to the
stability of a native protein.
Figure 5.4 shows an ion pair within a folded protein. An anionic aspartic acid (charge = -1)
engages in attractive electrostatic interactions with cationic arginine (charge = +1). The dashed
lines represent hydrogen bonds.
Hydrogen bonding:
The idea that a single hydrogen atom could interact simultaneously with two other atoms was
proposed in 1920 by Latimer and Rodebush and their advisor, G. N. Lewis. Maurice Huggins,
who was also a student in Lewis' lab, describes the hydrogen bond in his 1919 dissertation. A
hydrogen bond is a favourable interaction between an atom with a basic lone pair of electrons
(a Lewis Base) and a hydrogen atom that has been partially stripped of its electrons because it
is covalently bound to an electronegative atom (N, O, or S). In a hydrogen bond, the Lewis
Base is the hydrogen bond acceptor (A) and the partially exposed proton is bound to the
hydrogen bond donor (H-D). Why hydrogen? Hydrogen is special because it is the only atom
that (i) forms covalent sigma bonds with electronegative atoms like N, O and S, and (ii) uses
the inner shell (1S) electron(s) in that covalent bond. When its electronegative bonding partner
pulls the bonding electrons away from hydrogen, the hydrogen nucleus (a proton) is exposed
on the back side (distal from the bonding partner). The unshielded face of the proton is
exposed, attracting the partial negative charge of an electron lone pair. Hydrogen is the only
atom that exposes its nucleus this way. Other atoms have inner shell non-bonding electrons
that shield the nucleus.
Figure 5.5 illustrates the elements of a hydrogen bond, including the HB acceptor and HB
donor, the lone pair and the exposed proton. N, O, S are the predominant hydrogen bonding
atoms (A & D) in biological systems.
A hydrogen bond is not an acid-base reaction, where the proton (H+) is fully transferred from
H-D to A to form D- and HA+. However, the strength of a hydrogen bond correlates well with
the acidity of donor H-D and the basicity of acceptor A. In a hydrogen bond, the H+ is partially
transferred from H-D to A, but H+ remains covalently attached to D. The H-D bond remains
intact.
Figure 5.6 Illustrates three different styles for representing a hydrogen bond. Atom A is the
Lewis base (for example the N in NH3 or the O in H2O) and the atom D is electronegative (for
example O, N or S). The conventional nomenclature is confusing: a hydrogen bond is not a
covalent bond.
Figure 5.7 shows the most common hydrogen bond acceptors and donors in biological
macromolecules.
The most common hydrogen bonds in biological systems involve oxygen and nitrogen atoms
as A and D. Keto groups (=O), amines (R3N), imines (R=N-R) and hydroxyl groups (-OH) are
the most common hydrogen bond acceptors in DNA, RNA, proteins and complex
carbohydrates. Hydroxyl groups and amines/imines are the most common hydrogen bond
donors. Hydroxyls and amines/imines can both donate and accept hydrogen bonds.
In traversing the Period Table, increasing the electronegativity of atom D strips electron density
from the proton (in H-D), increasing its partial positive charge, and increasing the strength of
any hydrogen bond. Thiols (-SH) can both donate and accept hydrogen bonds but these are
generally weak, because sulfur is not sufficiently electronegative. Hydrogen bonds involving
carbon, where H-D equals H-C, are observed, although these are weak and infrequent.
C is insufficiently electronegative to form good hydrogen bonds. Hydrogen bonds are
essentially electrostatic in nature, although the energy can be decomposed into additional
contributions from polarization, exchange repulsion, charge transfer, and mixing.
Hydrogen bond strengths form a continuum. Strong hydrogen bonds of 20-40 kcal/mole (82
to 164 kjoule/mole), generally formed between charged donors and acceptors, are nearly as
strong as covalent bonds, Weak hydrogen bonds of 1-5 kcal/mole (4 - 21 kjoule/mole),
sometimes formed with carbon as the proton donor, are no stronger than conventional dipole-
dipole interactions. Moderate hydrogen bonds, which are the most common, are formed
between neutral donors and acceptors are from 3 - 12 kcal/mole (12 - 50 kjoule/mole)). A
hydrogen bond is not a bond. It is a molecular interaction (a non-bonding interaction).
Figure 5.8 shows cooperativity of the hydrogen bonds of an acetic acid dimer (top) and of a
G-C base pair (bottom). One hydrogen bond increases the stability of the adjacent hydrogen
bond (and vice versa).
For example, in the hydrogen-bonded systems above (the acetic acid dimer), the top hydrogen
bond increases both the acidity of the hydrogen, and the basicity of the oxygen in the bottom
hydrogen bond. Each hydrogen bond makes the other stronger than it would be in isolation.
Cooperativity of hydrogen bonding is observed in base pairing and in folded proteins.
Figure 5.9 shows cooperativity via resonance of the hydrogen bonds of an anti-parallel β-sheet.
One property of molecules appears to be very close to a binary relation: that is two atoms in a
given molecule are either bonded or not bonded. Therefore, molecules can be represented by
graphs when the only property considered is the existence or not of a chemical bond. This
property is called molecular topology. We define molecular topology as the totality of
information contained in the molecular graph. In chemistry graphs can represent different
chemical objects: molecules, reactions, crystals, polymers, clusters, etc. The common feature
of chemical systems is the presence of sites and connections between them. Sites may be
atoms, electrons, molecules, molecular fragments, groups of atoms, intermediates, orbitals, etc.
The connections between sites may represent bonds of any kind, bonded and nonbonded
interactions, elementary reaction steps, rearrangements, van der Waals forces, etc. Chemical
systems may be depicted by chemical graphs using a simple conversion rule:
Site ↔ vertex
connection ↔ edge
A special class of chemical graphs are molecular graphs. Molecular graphs are chemical
graphs which represent the constitution of molecules. They are also called constitutional
graphs. In these graphs vertices correspond to individual atoms and edges to chemical bonds
between them. Molecular graphs are necessarily connected graphs. As examples the molecular
graphs corresponding to propane and cyclopropane are shown in Figure 5.10.
Figure 5.11 The hydrogen suppressed molecular graphs depicting butane and cyclobutene
The molecular graph grossly simplifies the complex picture of a molecule by depicting only
its constitution (i.e., the chemical bonds between the various pairs of atoms in the molecule)
and neglecting other structural features (e.g., geometry, stereochemistry, chirality). Even so, a
simple picture of a molecule as the molecular graph can enable one to make useful predictions
about physical and chemical properties of molecules. Since the predictions of properties and
reactivities of molecules are of prime interest to chemists, the development of chemical graph
theory is, thus, justified.
Molecular graphs depicting constitutional formulae of molecules represent their topology. This
is a chemist’s view of molecular topology. However, a more precise definition of molecular
topology may also be given using the concept of the molecular graph. A topological space is
formed by a set and the topological structure defined upon the set. A simple connected
(molecular) graph can be associated with a topological space if it can be shown that a
topological structure is defined upon its vertex-set.
Graphs, adequately labeled, may be associated with several matrices. A graph G is labeled if
a certain numbering of vertices of G is introduced. Here two graph-theoretical matrices, i.e.,
the adjacency matrix and the distance matrix will be discussed. They are also sometimes
referred to as topological matrices. These matrices may be used for identifying certain
properties of graphs, which would not otherwise easily emerge.
The most important matrix representation of a graph G is the vertex- adjacency matrix A =
A(G). This matrix is also of importance in chemistry and physics. The vertex-adjacency matrix A(G)
of a labeled connected graph G with N vertices is the square N x N symmetric matrix which contains
information about the internal connectivity of vertices in G. It is defined as,
Aii = 0 -------(5.2)
Therefore, a nonzero entry appears in A(G) only if an edge connects vertices i and j. For
example, the following vertex-adjacency matrix can be constructed for a labeled graph
G (Figure 5.12).
Figure 5.12 A vertex and edge labelled graph G
The adjacency matrix is symmetrical about the principal diagonal. There- fore, the
transpose of the adjacency matrix A leaves the adjacency matrix unchanged
For example, the following edge-adjacency matrix can be constructed for a labeled graph G
in Figure 5.12
Although both the vertex-adjacency matrix and the edge-adjacency matrix reflect the topology
of a molecule, they differ in their structure. However, it should be noted while the vertex-
adjacency matrix uniquely determines the graph, the edge-adjacency matrix does not. In other
words, there are known non-isomorphic graphs with identical edge-adjacency matrices. A pair
of such non-isomorphic graphs is shown in Figure 5.13. The corresponding edge-adjacency
matrix is given by
Figure5.13 A pair of nonisomorphic graphs (G1 and G2) which possess the identical edge-
adjacency matrix.
The distance matrix (which is also sometimes called the metrics matrix) is, in a sense, a more
complicated and also a richer structure than the adjacency matrix. It is a graph-theoretical
(topological) matrix less common than the adjacency matrix, but it has been increasingly used
in the last two decades. in many different areas of chemistry and physics. It has been pointed
out an interesting fact that the distance matrix has also found considerable use in the areas of
research which are relatively remote from chemistry and physics and to a great extent non
mathematical, such as anthropology, geography, geology, ornithology, philology, and
psychology.
The distance matrix D = D(G) of a labeled connected graph G is a real symmetric N x N matrix
whose elements (D)ij are defined as follows:
where 𝑒𝑖𝑗,is the length of the shortest path (i.e., the minimum number of edges) between the
vertices vi , and vj. The length is also called the distance between the vertices vi, and vj thence
the term distance matrix. For example, the following distance matrix can be constructed for
a labeled graph G (Figure 5.15):
Figure 5.15 A label graph G and its distance matrix
The distance matrix has found a widespread application in chemistry in both explicit and
implicit forms. The first explicit use of the distance matrix was employed the for studying the
permutational isomers of stereo chemically nonrigid molecules. The distance matrix in explicit
form is also used to generate the distance polynomial and the distance spectrum.
More examples of adjacency (vertex, edge) matrix and distance matrix are given below in
Table 5.1.
Table 5.1 Examples of adjacency and distance matrix for few graphs
Graphs Vertex adjacency Edge adjacency Distance
5.5 Topological indices
Topological indices, introduced over 150 years ago, remain widely used due to their durability
and versatility, with more than 120 reported to date. However, their rapid growth highlights
the lack of clear criteria for selection and verification, as many indices are strongly
intercorrelated and often encode similar structural information. Most are derived from the
adjacency or distance matrix of molecular graphs, the latter itself originating from the
adjacency matrix. Despite their redundancy, topological indices serve as powerful molecular
descriptors, supporting the long-term goal of predicting bulk properties of matter directly from
molecular structure. Topological indices are numerical parameters derived from the graph
representation of a molecule, where atoms are treated as vertices and bonds as edges. They
capture structural information such as size, shape, branching, and connectivity of the molecule,
and are widely used as molecular descriptors in QSAR/QSPR studies to correlate chemical
structure with physical, chemical, or biological properties.
A single number that can be used to characterize the graph of a molecule is called a topological
index. (The term graph-theoretical index would be more accurate than topological index, but
the latter is more common in the chemical literature.) A topological index, thus, appears to be
a convenient device for converting chemical constitution into a number. Evidently, this number
must have the same value for a given molecule regardless of ways in which the corresponding
graph is drawn or labeled. Such a number is referred to by graph theorists as a graph invariant.
For example, one of the simplest graph invariants (topological indices) is the number of
vertices in the graph (the number of atoms in the molecule). Hence, it could be simply said
that topological indices are graph invariants. It should also be pointed out that topological
indices do not generally allow the reconstruction of the molecular graph, implying that a
certain loss of information has occurred during their creation.
The interest in topological indices is in the main related to their use in nonempirical
quantitative structure-property relationships (QSPR) and quantitative structure-activity
relationships (QSAR). They are used to predict biological activity, toxicity, and
physicochemical properties such as boiling point, solubility, and stability. They are also
applied in molecular similarity analysis, classification of compounds, materials design, and
environmental chemistry. In simple terms, topological indices provide a way to translate
molecular structure into numerical values that can be used for property prediction and
comparison across molecules.
Zagreb indices:
The first and second Zagreb indices (M1 and M2) are another set of classic vertex-based
descriptors developed in 1972 and 1975, respectively. They were called the Zagreb group
indices as their authors were members of the “Rudjer Bošković” Institute in Zagreb, Croatia.
In these indices one counts the connections from each vertex (node, carbon). The first Zagreb
index M 1(G) is equal to the sum of squares of the degrees of the vertices, and the second
Zagreb index M 2(G) is equal to the sum of the products of the degrees of pairs of adjacent
vertices of the underlying molecular graph G.
-------(5.7)
There are thousands of 2D descriptors that are frequently applied in modeling or predicting
properties or biological functions. What is interesting is that these graphs are often descriptors
that are reduced to a single value that can be used to make meaning of the physical world.
Zagreb group indices were introduced to characterize branching.
Wiener index:
One of the first mathematical representations of chemical structure used for prediction of
properties was developed in 1947 by Harold Weiner. It is defined at the sum of distances
between any two carbon atoms (pairs of nodes) in the molecule. Mathematically it is
represented as:
-------(5.8)
To calculate the Wiener index for a molecule, for each pair of atoms in the structure, count the
distance between atoms. Take the sum of all distances and divide by two. For example in the
case of ethane, which only has two nodes:
u 0 1
v 1 0
Pentane has 5 nodes, and distances between each node are calculated and summed.
A B C D E total
A 0 1 2 3 4 10
B 1 0 1 2 3 7
C 2 1 0 1 2 6
D 3 2 1 0 1 7
E 4 3 2 1 0 10
Platt was also interested in devising a scheme for predicting physical parameters (molar
volumes, boiling points, heats of formation, heats of vaporization) of alkanes. He introduced
an index F = F(G), which is equal to the total sum of edge-degrees in a graph G. The edge-
degree of an edge e, D(e), is the number of its adjacent edges. This index was named the Platt
number. The Platt number of a graph G is defined by
F(G) =∑𝑀
𝑖=1 𝐷 (𝑒𝑖) -------(5.9)
Few examples of calculated Zagreb index, Wiener index and Platt number for few graphs is
given in Table 5.2
[1, 2, 1] 6 4 4 2
[1, 2, 2, 1] 10 8 10 4
[1, 2, 2, 2,
1] 14 12 20 6
[2, 2, 2] 12 12 3 6
[2, 2, 2, 2] 16 16 8 8
[2, 2, 2, 2,
2] 20 20 15 10
[1, 1, 1, 3] 12 9 9 6
[3, 3, 3, 3] 36 54 6 24
QSPR:
The first step in developing a QSPR equation is to compile a list of compounds for which the
experimentally determined property is known. Ideally, this list should be very large. Often,
thousands of compounds are used in a QSPR study. If there are fewer compounds on the list
than parameters to be fitted in the equation, then the curve fit will fail. If the same number
exists for both, then an exact fit will be obtained. This exact fit is misleading because it fits the
equation to all the anomalies in the data, it does not necessarily reflect all the correct trends
necessary for a predictive method. In order to ensure that the method will be predictive, there
should ideally be 10 times as many test compounds as fitted parameters. The choice of
compounds is also important. For example, if the equation is only fitted with hydrocarbon data,
it will only be reliable for predicting hydrocarbon properties. The next step is to obtain
geometries for the molecules. Crystal structure geometries can be used; however, it is better to
use theoretically optimized geometries. By using the theoretical geometries, any systematic
errors in the computation will cancel out. Furthermore, the method will predict as yet un-
synthesized compounds using theoretical geometries. Some of the simpler methods require
connectivity only.
Molecular descriptors must then be computed. Any numerical value that describes the
molecule could be used. Many descriptors are obtained from molecular mechanics or
semiempirical calculations. Energies, population analysis, and vibrational frequency analysis
with its associated thermodynamic quantities are often obtained this way. Ab initio results can
be used reliably, but are often avoided due to the large amount of computation necessary. The
largest percentage of descriptors are easily determined values, such as molecular weights,
topological indexes, moments of inertia, and so on. Table 30.1 lists some of the descriptors
that have been found to be useful in previous studies. These are discussed in more detail in the
review articles listed in the bibliography.
Once the descriptors have been computed, is necessary to decide which ones will be used. This
is usually done by computing correlation coefficients. Correlation coefficients are a measure
of how closely two values (descriptor and property) are related to one another by a linear
relationship. If a descriptor has a correlation coefficient of 1, it describes the property exactly.
A correlation coefficient of zero means the descriptor has no relevance. The descriptors with
the largest correlation coefficients are used in the curve fit to create a property prediction
equation. There is no rigorous way to determine how large a correlation coefficient is
acceptable.
Intercorrelation coefficients are then computed. These tell when one descriptor is redundant
with another. Using redundant descriptors increases the amount of fitting work to be done,
does not improve the results, and results in unstable fitting calculations that can fail completely
(due to dividing by zero or some other mathematical error). Usually, the descriptor with the
lowest correlation coefficient is discarded from a pair of redundant descriptors.
A curve fit is then done to create a linear equation, such as Property = c0 + c1d1 + c2d2 + ···,
where ci are the fitted parameters and di the descriptors. Most often, the equation being fitted
is a linear equation like the one above. This is because the use of correlation coefficients and
linear equations together is an easily automated process. Introductory descriptions cite linear
regression as the algorithm for determining coefficients of best fit, but the mathematically
equivalent matrix least- squares method is actually more efficient and easier to implement.
Occasionally, a nonlinear parameter, such as the square root or log of a quantity, is used. This
is done when a researcher is aware of such nonlinear relationships in advance.
[Start]
↓
[1. Compile List of Compounds]
- Use compounds with known experimental property
- Prefer large datasets
- Ensure chemical diversity (not just hydrocarbons)
↓
[2. Obtain Molecular Geometries/topology]
→ Use theoretical geometries (preferred over crystal structures)
→ For simple methods, connectivity may suffice
↓
[3. Compute Molecular Descriptors]
- From molecular mechanics / semiempirical methods
- Include energies, population analysis, vibrational frequencies, etc.
- Easily obtained descriptors: molecular weight, topological index, etc.
↓
[4. Select Descriptors Based on Correlation Coefficients]
- Calculate correlation between each descriptor and property
- Select descriptors with highest correlation
↓
[5. Remove Redundant Descriptors (Intercorrelation)]
- Compute intercorrelation between selected descriptors
- Discard descriptor with lower property correlation if redundant
↓
[6. Perform Curve Fitting]
- Fit linear equation: Property = c₀ + c₁d₁ + c₂d₂ + ...
- Use linear regression or matrix least-squares
- Occasionally apply nonlinear terms (e.g., √, log) if justified
↓
[7. Generate Predictive Equation]
- Use selected descriptors and fitted parameters
↓
[8. Apply for Property Prediction]
→ Use equation to predict new or un-synthesized compounds
↓
[End]
QSAR:
The concept of Quantitative Structure–Activity Relationship (QSAR) is central to in silico
prediction of chemical, biological, and physical properties of molecules. QSAR is based on the
idea that the biological activity or physicochemical property of a compound is directly related
to its chemical structure, and therefore can be predicted mathematically. In practice, molecular
structures are translated into descriptors—numerical values that capture features such as
hydrophobicity, electronic distribution, steric effects, or topological indices. These descriptors
are then correlated with experimentally known activities or properties using statistical or
machine learning models. Once a reliable QSAR model is developed, it can be applied to
predict the properties of new, untested compounds without the need for costly and time-
consuming laboratory experiments. This approach is widely used in drug discovery for
predicting potency, toxicity, and pharmacokinetics, as well as in environmental chemistry for
assessing biodegradability and pollutant toxicity. The efficiency of QSAR lies in its ability to
screen large molecular libraries computationally and prioritize the most promising candidates
for synthesis and experimental validation, thus saving both time and resources.
QSAR is also called traditional QSAR or Hansch QSAR to distinguish it from the 3D QSAR
method. This is the application of the technique described above to biological activities, such
as environmental toxicology or drug activity. The discussion above is applicable but a number
of other caveats apply; which are addressed in this section. The following discussion is oriented
toward drug design, although the same points may be applicable to other areas of research as
well.
In order to parameterize a QSAR equation, a quantified activity for a set of compounds must
be known. These are called lead compounds, at least in the pharmaceutical industry. Typically,
test results are available for only a small number of compounds. Because of this, it can be
difficult to choose a number of descriptors that will give useful results without fitting to
anomalies in the test set. Three to five lead compounds per descriptor in the QSAR equation
are normally considered an adequate number. If two descriptors are nearly collinear with
one another, then one should be omitted even though it may have a large correlation
coefficient.
In the case of drug design, it may be desirable to use parabolic functions in place of linear
functions. The descriptor for an ideal drug candidate often has an optimum value. Drug activity
will decrease when the value is either larger or smaller than optimum. This functional form is
described by a parabola, not a linear relationship. The advantage of using QSAR over other
modeling techniques is that it takes into account the full complexity of the biological system
without re- quiring any information about the binding site. The disadvantage is that the method
will not distinguish between the contribution of binding and trans- port properties in
determining drug activity. QSAR is very useful for deter- mining general criteria for activity,
but it does not readily yield detailed structural predictions.
Computational chemistry employs molecular modelling to study and predict the structure,
properties, and interactions of molecules. It offers wide scope in research and industry due to
its cost-effectiveness and efficiency compared to experimental methods. Molecular interactions
can be analyzed using mathematical and topological approaches. Molecular topology provides
a systematic way to represent molecules using matrices such as vertex adjacency, edge
adjacency, and distance matrices. From these, topological indices like the Zagreb index, Wiener
index and Platt number are derived to quantify structural features. These descriptors are further
applied in QSPR (Quantitative Structure–Property Relationships) and QSAR (Quantitative
Structure–Activity Relationships) studies, which enable prediction of molecular properties and
biological activities, thereby aiding in drug design, material development and environmental
chemistry.
References
[2] C. J. Cramer, Essentials of Computational Chemistry: Theories and Models, 2nd ed.
Hoboken, NJ, USA: John Wiley & Sons, 2023, ISBN: 978-0-470-09182-1.
[3] F. Jensen, Introduction to Computational Chemistry, 3rd ed. Hoboken, NJ, USA: Wiley,
2017.
[5] N. Trinajstić, Chemical Graph Theory. Boca Raton, FL, USA: CRC Press, 1992.