Machine Learning in Material Innovation
Machine Learning in Material Innovation
Review
Application of Machine Learning in Material Synthesis and
Property Prediction
Guannan Huang, Yani Guo, Ye Chen and Zhengwei Nie *
School of Mechanical and Power Engineering, Nanjing Tech University, Nanjing 211816, China;
i35653184@[Link] (G.H.); guo17793470078@[Link] (Y.G.); chenye@[Link] (Y.C.)
* Correspondence: niez@[Link]
Abstract: Material innovation plays a very important role in technological progress and industrial
development. Traditional experimental exploration and numerical simulation often require con-
siderable time and resources. A new approach is urgently needed to accelerate the discovery and
exploration of new materials. Machine learning can greatly reduce computational costs, shorten the
development cycle, and improve computational accuracy. It has become one of the most promising
research approaches in the process of novel material screening and material property prediction.
In recent years, machine learning has been widely used in many fields of research, such as super-
conductivity, thermoelectrics, photovoltaics, catalysis, and high-entropy alloys. In this review, the
basic principles of machine learning are briefly outlined. Several commonly used algorithms in
machine learning models and their primary applications are then introduced. The research progress
of machine learning in predicting material properties and guiding material synthesis is discussed.
Finally, a future outlook on machine learning in the materials science field is presented.
1. Introduction
Citation: Huang, G.; Guo, Y.; Chen,
New materials have become the cornerstone of scientific and technological develop-
Y.; Nie, Z. Application of Machine
ment. Discovering materials with targeted properties, especially nanomaterials, has always
Learning in Material Synthesis and
Property Prediction. Materials 2023,
been a hotspot in science [1,2]. At present, the research and development of new materials
16, 5977. [Link]
mainly relies on researchers’ intuitive judgment of materials and empirical trial-and-error
10.3390/ma16175977
methods, which are not only inefficient but also often require a certain level of experience
and luck to obtain the target materials. At the same time, methods based on density func-
Academic Editor: Dorota
tional theory (DFT) are widely used in the research and development of novel materials.
Wilk-Kołodziejczyk
Since their initial development, DFT methods have evolved from limited calculations that
Received: 31 July 2023 provide approximate results to increasingly accurate and predictable methods. These meth-
Revised: 22 August 2023 ods have made important contributions in a variety of fields, such as materials discovery
Accepted: 28 August 2023 and design, drug design, solar cells, and hydrolytic materials [3]. The accuracy of these
Published: 31 August 2023 methods, however, is limited in fast calculations. To obtain high-accuracy results, the
computational volume often has to be much higher, which is difficult to exploit efficiently
in the research and development of new materials. In this context, artificial intelligence (AI)
is becoming highly popular with researchers as a means of accelerating the development of
Copyright: © 2023 by the authors. innovative materials. A subfield of AI that has grown rapidly in recent years is machine
Licensee MDPI, Basel, Switzerland.
learning (ML). ML applications are built on statistical algorithms. ML performs similarly
This article is an open access article
to researchers’ performance [4]. Because of its powerful data processing capability and
distributed under the terms and
relatively low research threshold, ML can effectively reduce human and material costs in
conditions of the Creative Commons
the process of novel material development and shorten the research and development cycle.
Attribution (CC BY) license (https://
By replacing or collaborating with traditional experiments and computational simulations,
[Link]/licenses/by/
4.0/).
ML could be employed to analyze material structures and predict material properties,
predict material properties, enabling the development of novel functional materials more
enabling the development of novel functional materials more efficiently and accurately.
efficiently and accurately. As a result, ML has become one of the most crucial methods for
As a result, ML has become one of the most crucial methods for replacing traditional
replacing traditional research and development. In the recent past, researchers in different
research and development. In the recent past, researchers in different fields, including
fields, including computer scientists and experts in AI algorithms, have used this ap-
computer scientists and experts in AI algorithms, have used this approach extensively,
proach contributing
greatly extensively, to
greatly contributingoftoML
the development thetechniques
development of ML
[5]. ML techniques
is now [5]. ML in
widely utilized is
now widely utilized in fields such as natural language understanding,
fields such as natural language understanding, non-monotonic reasoning, machine vision, non-monotonic
reasoning,
and patternmachine vision,
recognition [6]. and pattern recognition [6].
The basic
The basic principle
principle of ML is
of ML is to
to learn
learn (or
(or guess)
guess) general
general patterns
patterns from
from aa limited
limited
amount of training data and use these patterns to make predictions on unknown data.
amount of training data and use these patterns to make predictions on unknown data.
Figure 11 shows
Figure showsan anML
MLworkflow
workflow example.
example. MLML hashasbeenbeen
usedused
to detect the solubility
to detect of C60
the solubility
in materials
of science as
C60 in materials early as
science as the lastascentury
early the last[7]. It is now
century [7].used
It isto
nowdiscover
used novel mate-
to discover
rials, predict material and molecular properties, study quantum
novel materials, predict material and molecular properties, study quantum chemistry,chemistry, and design
drugs.
and The drugs.
design purposeTheof purpose
this review is toreview
of this offer an overview
is to offer an of the employment
overview of ML in
of the employment
predicting material properties and performance, guiding material
of ML in predicting material properties and performance, guiding material synthesis, synthesis, and project-
ing models and conclusions. This review not only provides guidance
and projecting models and conclusions. This review not only provides guidance for for researchers to
synthesize stable and efficient materials, but also inspires their interest in
researchers to synthesize stable and efficient materials, but also inspires their interest in the use of ML
in materials
the use of MLresearch.
in materials research.
workflow.
Figure 1. An example of an ML workflow.
2. Data Pre-Processing
2. Data Pre-Processing
If
If ML models
ML models are
are the
the engines
engines that
that handle
handle various
various tasks,
tasks, data
data are
are the
the fuel
fuel that
that drives
drives
the models. A sufficient amount of data is a prerequisite to making the model
the models. A sufficient amount of data is a prerequisite to making the model work. High- work.
High-quality
quality data enable the model to run effectively. Due to this, large amounts of datadata
data enable the model to run effectively. Due to this, large amounts of are
are critical
critical to ML
to ML [8]. [8]. In general,
In general, the final
the final ML ML results
results are directly
are directly affected
affected by amount
by the the amount
and
and reliability of the data. This is where data pre-processing and feature engineering are
reliability of the data. This is where data pre-processing and feature engineering are ben-
beneficial. Data pre-processing and feature engineering could promote the reconstruction of
eficial. Data pre-processing and feature engineering could promote the reconstruction of
datasets so that computers could more easily understand the physicochemical relationships
datasets so that computers could more easily understand the physicochemical relation-
of materials, detect material properties, and build prediction models [9].
ships of materials, detect material properties, and build prediction models [9].
2.1. Data Collection and Cleaning
2.1. Data Collection and Cleaning
2.1.1. Data Collection
[Link]
Data
ML,Collection
the size and quality of the training dataset employed for learning could
In ML, the
significantly sizethe
affect and quality of
accuracy of athe training model.
predictive dataset Therefore,
employed training
for learning couldneed
datasets sig-
nificantly
to affectorthe
be collected accuracy
created of a predictive
carefully. In general, model. Therefore,
training data cantraining datasets
be gathered needways.
in three to be
collected ordata
Obtaining created
fromcarefully. In general,
the published training
literature is thedata
firstcan be gathered
method. in three
The data ways.
obtained Ob-
in this
taining
way databefrom
could morethe published
relevant and literature
provide aisdirection
the first for
method. The data
synthesis obtained in [10].
and application this
way could
Second, be more relevant
high-throughput and provide or
computations a direction
experiments for synthesis andtoapplication
can be used obtain data.[10].
It
Second,behigh-throughput
should noted that, in some computations
cases, theseordata experiments can be used
may be incomplete, to obtain or
inconsistent, data.
even It
should be[11].
spurious notedThe that,third
in some cases,
method is these data may
to obtain data be incomplete,
from inconsistent,
open databases or even
available on
repository websites.
spurious [11]. The third The Materials
method is toGenome
obtain data Initiative,
from open initiated by the
databases Unitedon
available States
repos-in
2011,
itory emphasizes
websites. The theMaterials
importance of massive
Genome data in
Initiative, the development
initiated of materials
by the United science,
States in 2011,
which encourages
emphasizes the development
the importance of massive of high-quality material databases
data in the development [12]. With
of materials the
science,
continuous development of theoretical and experimental research,
which encourages the development of high-quality material databases [12]. With the con- data generated from
experiments and computational
tinuous development simulations,
of theoretical including failure
and experimental data,data
research, have generated
been integrated
from
Materials 2023, 16, 5977 3 of 30
into databases [13]. These databases are based on the concept of material data sharing,
which greatly simplifies the process of obtaining material information. Table 1 introduces
some commonly used methods for collecting data from publicly available databases.
For instance, Zhou et al. [14] developed an ML-based approach to predict cathode
materials for Zn-ion batteries with high capacity and high voltage. They screened over
130,000 inorganic materials from the materials project database and applied a crystal
graph convolutional-neural-network-based ML approach with data from the Automatic
Flow (AFLOW) database. This resulted in the prediction of approximately 80 cathode
materials, with 10 of them being experimentally discovered previously and agreeing
well with the observed measurements. Additionally, approximately 70 new promising
candidates were predicted for further experimental validation.
Figure 2. Evolution of the ML workflow in nanomaterial discovery and design. (a) First-generation
approach. In this paradigm, there are two main steps: feature engineering from raw database to de-
scriptors and model building from descriptors to target model. (b) Second-generation approach. The
key characteristic that distinguishes this approach from the first-generation approach is eliminating
human-expert feature engineering, which can directly learn from raw nanomaterials. Reproduced
with permission from [21].
Materials 2023, 16, 5977 5 of 30
3.1.1. KNN
The KNN algorithm was first proposed by Cover and Hart [25]. The KNN classification
is one of the most basic and simplest classification methods. It should be considered for
classification studies when little or no data distribution experience is available [26]. The
principle of the KNN algorithm is that if most of the most similar K samples in the feature
space (i.e., the nearest samples in the feature space) belong to a certain category, the sample
also belongs to this category. Figure 3 shows a schematic of a typical KNN algorithm. For
an unknown target, when K takes 3, the target is classified into class 1; when K takes 7,
the target is classified into class 2. According to this method, the sample’s category is
determined by its proximity to one or more nearby samples. The KNN algorithm itself
is simple and effective, easy to understand, and straightforward to implement. Since it
does not require prediction parameters or training, the KNN algorithm is suitable for time
classifications, especially for multimodals (i.e., objects with multiple categories). Recently,
KNN algorithms have been widely utilized in text classification, pattern recognition, image
processing, and materials science. Sharma et al. [27] employed the KNN algorithm to
predict the dynamic fracture toughness of glass-filled polymer composites. The dynamic
modulus of elasticity, aspect ratio, and volume fraction of glass particles were used as
independent model parameters. The proposed KNN model predicted the fracture behavior
of the composites with an accuracy of 96%. It is also possible to extend their model to
predict other material properties.
The drawback of the KNN algorithm is that as the amount of data increase, the
computational complexity of the KNN increases accordingly. This is because the KNN
algorithm needs to calculate both training data and test data for each classification or
regression. If there are a large amount of data, the computing power required would be
greatly increased. In addition, the randomness of training data also affects the performance
of the KNN algorithm [28].
Materials 2023, 16, x FOR PEER REVIEW 6 o
Materials 2023, 16, 5977 6 of 30
The drawback of the KNN algorithm is that as the amount of data increase, the c
putational complexity of the KNN increases accordingly. This is because the KNN a
rithm needs to calculate both training data and test data for each classification or reg
sion. If there are a large amount of data, the computing power required would be gr
increased. In addition, the randomness of training data also affects the performance o
KNN algorithm [28].
3.1.2. DT
A DT is a typical classification method. The earliest DT algorithm was the con
Figure 3. Schematic of a typical KNN algorithm.
Figure 3. Schematic
learning of a typicalby
system proposed KNN algorithm.
Hunt [29]. The most influential DT algorithms are ID3
3.1.2. DT
and C4.5 [31], which were proposed by Quinlan in 1986 and 1993, respectively. DTs
The
sify A DTdrawback
is a typical
training data by of the KNN algorithm
classification
different method.
features,The isearliest
that as
aiming the
toDT amountcategorize
algorithm
correctly of data
was increase,
the concept the A
instances. c
learning
putational system proposed
complexity by Hunt [29].
of thedecision The
KNN increases most influential
accordingly. DT algorithms
This are
is because ID3 [30]
model consists of internal nodes and leaf nodes. Each internal the nodeKNNsplita
and
rithmC4.5
needs[31], which were
to calculate proposed by Quinlan in 1986 and 1993, respectively. DTs
instance space into two both or moretraining data and
subspaces test data
according to for each classification
a certain
classify training data by different features, aiming to correctly categorize instances. A
discrete function or rego
sion.
input If there
attribute are a large amount of data, the computing power required would be gre
DT model consistsvalues, anddecision
of internal each leaf node
nodes and isleaf
assigned
nodes. toEachone class representing
internal node splits the
increased.
appropriate
the In
instance space addition,
target
intovalue the randomness
two or[32].
moreChen of training
et [Link]
subspaces data
[11] presented also affects
thediscrete
to a certain structurethe performance
of a typical
function of of
D
KNN
shown
the algorithm
input Figure[28].
inattribute 4. A typical
values, and eachdecision treeis algorithm
leaf node assigned toconsists
one classof three mainthe
representing steps: fea
most appropriate target value [32]. Chen et al. [11] presented
selection, decision tree generation, and pruning. The purpose of pruning is to minithe structure of a typical
DT, asDT
3.1.2. shown in Figure 4. A typical decision tree algorithm consists of three main steps:
the structural risk of the model by optimizing the loss function and weighing the mo
feature selection, decision tree generation, and pruning. The purpose of pruning is to
A
minimize DT
complexity is
the anda typical
accuracy.
structural classification
risk ofLiu
the et method.
al. [33]
model developed
by optimizingThe earliest
a DT
the DT algorithm
lossmodel
function forand was the
predicting
weighing the con
resi
learning
tensile system
strength proposed
and modulus by Hunt
of [29]. The most influential
pultruded-fiber-reinforced
the model’s complexity and accuracy. Liu et al. [33] developed a DT model for predicting DT
polymeralgorithms
(FRP) are ID3
compo
and
Using
the C4.5an[31],
residual which
existing
tensile wereand
database,
strength proposed
746
modulusdataby Quinlan
points
of were in collected
1986 and for
pultruded-fiber-reinforced 1993, respectively.
training.
polymer DTs c
The accurac
(FRP)
composites.
sify trainingwas
the model Using
data an existing
by different
verified database, 746 data
features, The
experimentally. points
aiming were collected
to correctly
significance for training.
of allcategorize The
attributes instances.
of the input A
accuracy
model
was also ofquantitatively
the model
consists was analyzed
of internal verified
decisionexperimentally.
bynodes andThe
the model. leaf
Thesignificance
nodes.
proposed ofDT
Each allinternal
attributes
model node of splits
provides a
the input data was also quantitatively analyzed by the model. The proposed DT model
instance
method space into twothe
for predicting or long-term
more subspaces according
degradation of FRPto acomposites
certain discrete subjectedfunction
to envof
provides a new method for predicting the long-term degradation of FRP composites
input
mental attribute
subjected influences. values, and
to environmental each leaf node is assigned to one class representing the m
influences.
appropriate target value [32]. Chen et al. [11] presented the structure of a typical DT
shown in Figure 4. A typical decision tree algorithm consists of three main steps: fea
selection, decision tree generation, and pruning. The purpose of pruning is to minim
the structural risk of the model by optimizing the loss function and weighing the mod
complexity and accuracy. Liu et al. [33] developed a DT model for predicting the resid
tensile strength and modulus of pultruded-fiber-reinforced polymer (FRP) compos
Using an existing database, 746 data points were collected for training. The accurac
the model was verified experimentally. The significance of all attributes of the input d
was also quantitatively analyzed by the model. The proposed DT model provides a n
method for predicting the long-term degradation of FRP composites subjected to envi
mental influences.
Figure 4. Diagram of a DT. The circles and squares indicate internal nodes and leaf nodes, respectively.
Different colors represent different classes. Reproduced with permission from [11].
Figure 4. Diagram of a DT. The circles and squares indicate internal nodes and leaf nodes, res
tively. Different colors represent different classes. Reproduced with permission from [11].
The RF algorithm consists of multiple DTs. In RFs, each tree casts a unit vote for
Materials 2023, 16, 5977 7 of 30
most popular class, and then combining these votes obtains the final sort result. RFs
sess high classification accuracy [34]. It would, however, take a great deal of space
timeThe
to RF
train an RF consists
algorithm with many DTs. Compared
of multiple DTs. In RFs,with
each DTs, the acalculation
tree casts unit vote forcosts of
would also increase significantly. In this regard, RFs and DTs should be selected
the most popular class, and then combining these votes obtains the final sort result. RFs base
the actual
possess situation.
high classification accuracy [34]. It would, however, take a great deal of space
and time to train an RF with many DTs. Compared with DTs, the calculation costs of RFs
would also increase significantly. In this regard, RFs and DTs should be selected based
3.1.3. ANN
on the actual situation.
The concept of an ANN was introduced by McCulloch and Pitts [35]. An ANN
3.1.3. ANNnetwork structure that is formed by a large number of nodes (neurons) conne
complex
to each
Theother.
concept Itof
is an
a kind
ANNof abstraction,
was introduced simplification,
by McCulloch and and simulation
Pitts [35]. An ANN of the
is organiza
a
complex network structure that is formed by a large number of nodes
and operation mechanism of the human brain. Each node in an ANN represents a spe (neurons) connected
to each other. It is a kind of abstraction, simplification, and simulation of the organization
output function, i.e., the activation function. Each connection between any two nodes
and operation mechanism of the human brain. Each node in an ANN represents a specific
resents a weighted
output function, value
i.e., the for thefunction.
activation signal passing through that
Each connection connection,
between which is equ
any two nodes
lent to theamemory
represents weightedofvaluethe [Link] theThe network’s
signal connection
passing through that mode, the value
connection, whichofisthe weig
and the excitation
equivalent function
to the memory all have
of the [Link] Theeffect on its connection
network’s output [36]. mode,As athemajor soft-compu
value of
the weights, and
technology, ANNsthe excitation
have been function all have
extensively an effectand
studied on its outputin
applied [36]. As adecades
recent major [37].
soft-computing technology,
The structure ANNsANN
of a typical have been extensively
is shown studied
in Figure andnodes
5. Its appliedareingenerally
recent divi
decades [37].
into three categories: input, hidden, and output. The input nodes represent the in
The structure of a typical ANN is shown in Figure 5. Its nodes are generally divided
mation
into threereceived
categories:from
input,the inputand
hidden, data. TheThe
output. output
inputnodes are utilized
nodes represent to store the resul
the information
the datafrom
received processing.
the inputThe [Link] between
The output the are
nodes input and output
utilized to store nodes are of
the results so-called
the hid
nodes. DifferentThe
data processing. types
nodesof between
nodes inthe aninput
ANNand areoutput
distributed
nodes arein multiple layers. The no
so-called hidden
nodes.
on Different
different typescould
layers of nodes be in an ANN are
connected bydistributed
lines, whichin multiple
[Link] nodes in ne
synapses
on different layers could be connected by lines, which correspond
structures, representing a nonlinear mapping. The learning process of an ANN is to to synapses in neural
structures, representing a nonlinear mapping. The learning process of an ANN is to
tinuously optimize the whole network model by correcting the weights of nodes in e
continuously optimize the whole network model by correcting the weights of nodes in each
layer withtraining
layer with trainingdatadata
[38].[38].
Figure [Link]
Figure Diagramof aoftypical ANN.
a typical ANN.
A variety of ANN models and their variants have been developed. The variants
A variety of ANN models and their variants have been developed. The variant
include back-propagation networks, perceptrons, self-organizing mappings, Hopfield
clude back-propagation
networks, networks,ANNs
and Boltzmann machines. perceptrons,
have beenself-organizing
applied to drivemappings, Hopfield
the synthesis
works,
of a wide range of functional materials, such as shape memory alloys [39], hyperelastic of a w
and Boltzmann machines. ANNs have been applied to drive the synthesis
materials
range of [40], and high-entropy
functional materials,alloys
such(HEAs) [41]. memory
as shape Table 2 illustrates the application
alloys [39], of mate
hyperelastic
the afore-mentioned algorithms.
[40], and high-entropy alloys (HEAs) [41]. Table 2 illustrates the application of the af
mentioned algorithms.
Researchers
Materials 2023, 16, x FOR PEER REVIEW Algorithms Purposes 8 of 30
Predict the fracture toughness of silica-filled
Sharma et al. [42] KNN
epoxy composites.
Predict surface roughness in the micro-plasma
Researchers Algorithms Purposes
Kumar et al. [43] KNN transfer arc metal additive manufacturing
Sharma et al. [42] KNN Predict the fracture toughness (µ-PTAMAM)
of silica-filledprocess.
epoxy composites.
Predict surface roughness in the micro-plasma transfer arc metal
KumarJalali
et [Link][43]
al. [44] KNN KNN (Figure 6a) Predict phases in HEAs.
additive manufacturing (µ-PTAMAM) process.
Achieve rapid detection of transformer
Jalali Wang et al. [45]
et al. [44] KNN (Figure 6a) SVM Predict phases in HEAs.
winding materials.
Wang et al. [45] SVM Achieve rapid detection of transformer winding materials.
Predict the fracture life of martensitic steels under
Martinez et al. [46] SVM and ANN the fracture life of martensitic steels under high-tempera-
Predict
Martinez et al. [46] SVM and ANN high-temperature creep conditions.
ture creep conditions.
Adaptive boosting, RF, and DT Predict the compressive strength of concrete
Ahmad et al. [47] Adaptive boosting, RF,
Ahmad et al. [47] (Figure 6b) the compressive strengthatofhigh
Predict temperatures.
concrete at high temperatures.
and DT (Figure 6b)
Gradient boosted regression tree
Sun et al. [48] Gradient boosted regres- Evaluate the strength of coal–grout materials.
(GBRT) and RF
Sun et al. [48] Evaluate the strength of coal–grout materials.
sion tree (GBRT) and RF Predict the higher heating value (HHV) of biomass
Samadia et al. [49] GBRT
Predict the higher heating value (HHV)
materials based onof proximate
biomass materials
analysis. based
Samadia et al. [49] GBRT
on proximate
Predict analysis.
the compressive strength of eco-friendly
Shahmansouri et al. [50] Predict
ANN (Figure 6c)the compressive strength
geopolymer of eco-friendly
concrete incorporating geopolymer con-
silica fume and
Shahmansouri et al. [50] ANN (Figure 6c) natural zeolite.
crete incorporating silica fume and natural zeolite.
Development of a predictive
Development model for the chloride
of a predictive diffusion
model for coef-
the chloride
Liu etLiu
al. et al. [51]
[51] ANN ANN
diffusion coefficient in concrete.
ficient in concrete.
Figure 6. (a) A portion of the HEA interaction network with Fruchterman Reingold layout, adapted
with permission
with permission from
from [44].
[44]. (b)
(b) Schematic
Schematic illustration
illustration of
of an
an RF
RF structure,
structure, adapted
adapted with
with permission
permission
from [47]. (c) A multi-layer neural network model layout, adapted with permission from
from [47]. (c) A multi-layer neural network model layout, adapted with permission from [50].[50].
three recurrent neural networks (RNNs) to efficiently generate new energetic molecules
with high detonation velocity in the low data regime. They utilized data augmentation by
fragment shuffling of 303 energetic compounds to pretrain the RNN and then fine-tuned
it using the 303 compounds to produce molecules similar to the energetic compounds.
They also employed a simplified molecular input line entry (SMILE) system coupled with
pretrained knowledge to build an RNN-based prediction model for screening molecules
with high detonation velocity. Their strategy performed comparably to transfer learning
based on an existing big database. Quantum mechanics calculations confirmed that 35 new
molecules have higher detonation velocity and lower synthetic accessibility than the classic
explosive hexogen, with three novel molecules comparable to caged China Lake Compound
No. 20 in detonation velocity. Zhang et al. [65] utilized generative adversarial networks
(GANs) to design metaporous materials for sound absorption (Figure 7a). The researchers
trained the GANs using numerically prepared data and successfully developed designs
with high-standard broadband absorption performance. The GANs accelerated the design
process by hundreds of times, allowing for instantaneous multiple solutions. The GANs
also demonstrated the ability to generate creative configurations and rich local features.
This work highlighted the potential of ML in guiding the design and optimization process
for materials and opened up new possibilities for interdisciplinary research in AI and
materials. Unni et al. [66] introduced a deep convolutional mixture density network (MDN)
approach for the inverse design of layered photonic structures. The MDN modeled the
design parameters as multimodal probability distributions, allowing for convergence in
cases of nonuniqueness without sacrificing degenerate solutions. The MDN was applied to
the inverse design of two types of multilayer photonic structures consisting of thin films
of oxides, which present a challenge for conventional machine learning algorithms due
to their large degree of nonuniqueness in their optical properties. The MDN can handle
the transmission spectra of high complexity and varying illumination conditions. The
shape of the probability distributions provides valuable information for postprocessing
and prediction uncertainty. The MDN approach offers an effective solution to the inverse
design of photonic structures with high degeneracy and spectral complexity.
The use of vision transformers, residual networks (ResNets), and region-based-CNNs
(R-CNNs) on materials datasets has shown exceptional performance. Huang et al. [67]
proposed a waste materials classification method based on a vision transformer model
(Figure 7b). The model overcame CNN limitations by using self-attention mechanisms
to allocate weights to different parts of waste images. The vision transformer achieved
an accuracy rate of 96.98% by pretraining on ImageNet and fine-tuning on the TrashNet
dataset. The trained model can be deployed on a cloud server and accessed through
a portable device for real-time waste classification, which is convenient and efficient for
resource conservation and recycling. Jiang et al. [68] explored the use of global optimization
networks (GLOnets) with the ResNet architecture for the multiobjective and categorical
global optimization of photonic devices. The authors demonstrated that these networks,
called Res-GLOnets, could be configured to design thin-film stacks consisting of multiple
material types. The Res-GLOnets can find the global optimum with faster speeds compared
to conventional algorithms. The authors also showed the utility of their method for complex
design tasks, such as designing incandescent light filters. Wang et al. [69] proposed an
image detection method based on an improved Faster R-CNN model for wear location
and wear mechanism identification (Figure 7c). They trained and tested the model using
a wear image dataset produced by a self-made tribometer equipped with an imaging
system. The results showed that the proposed method had a detection accuracy of
more than 99%. It outperformed edge detection technology and Yolov3 target detection
models in wear location and wear mechanism identification. This research contributes
to the development of an innovative approach for the online and intelligent wear status
detection of machinery components.
Materials 2023, 16, x FOR PEER REVIEW 11 of 30
Materials 2023, 16, 5977 11 of 30
[Link]
Figure Some deep
deep learning
learning algorithm
algorithm structures.
structures. (a) Schematic
(a) Schematic illustration
illustration of procedures
of the design the design proce-
dures of metaporous materials with GANs, adapted with permission from [65].
of metaporous materials with GANs, adapted with permission from [65]. (b) Structure of a vision (b) Structure
trans- of a
vision transformer, adapted with permission from [67]. (c) Illustration of the concept of
former, adapted with permission from [67]. (c) Illustration of the concept of using image identification using image
identification
based based on
on the improved theR-CNN
Faster improved Faster
model R-CNN
to identify model
wear, to identify
adapted wear, adapted
with permission from [69].with permis-
sion from [69].
3.3. Materials Informatics Based on ML
3.3. Materials
Materials informatics
InformaticsisBased on ML
a study field that focuses on investigating and applying infor-
maticsMaterials
techniquesinformatics
to materialsisscience
a studyand engineering.
field Propelled
that focuses partly by theand
on investigating Materials
applying in-
Genome Initiative and partly by algorithmic developments and successes of
formatics techniques to materials science and engineering. Propelled partly by the Mate- data-driven
efforts in other domains,
rials Genome Initiative informatics strategies
and partly are beginning
by algorithmic to take shapeand
developments within materialsof data-
successes
science. Informatics strategies give rise to surrogate ML methods that can realize accurate
driven efforts in other domains, informatics strategies are beginning to take shape within
prediction using just historical data instead of experiments or simulations/calculations.
materials science. Informatics strategies give rise to surrogate ML methods that can realize
This methodology is usually composed of three distinct steps: acquisition of reliable
accurate data,
historical prediction using
statistical just historical
quantification data instead of material
of information-rich experiments or simulations/calcu-
structures, and map-
lations. This methodology is usually composed of three
ping between “input” and “output”. The commonly used ML algorithms in distinct steps: acquisition
materials of reli-
able historical
informatics data,
include statistical
regression, quantification
DT, ANN, and deep of learning
information-rich
[70–73]. Tomaterial
meet thestructures,
require- and
mapping
ments between
of the studies“input” and “output”.
of computational The informatics,
materials commonly used Zhao MLet [Link]
[74] derivedin an
materials
artificial-intelligence-aided data-driven infrastructure called Jilin Artificial-intelligence
informatics include regression, DT, ANN, and deep learning [70–73]. To meet the require-
aided
mentsMaterials-design
of the studies of Integrated Packagematerials
computational (JAMIP). informatics,
The organizationZhaoofetJAMIP abides
al. [74] derived an
by the data lifecycle in computational materials informatics, from data generation
artificial-intelligence-aided data-driven infrastructure called Jilin Artificial-intelligence to col-
aided Materials-design Integrated Package (JAMIP). The organization of JAMIP abides by
the data lifecycle in computational materials informatics, from data generation to collec-
tion and learning, as shown in Figure 8. It provides tools for materials production, high-
throughput calculations, data extraction and management, and ML-based data mining.
Materials 2023, 16, 5977 12 of 30
Materials 2023, 16, x FOR PEER REVIEW 12 of 30
lection and learning, as shown in Figure 8. It provides tools for materials production,
high-throughput calculations, data extraction and management, and ML-based data min-
The authors demonstrated the usefulness of JAMIP in exploring materials informatics in
ing. The authors demonstrated the usefulness of JAMIP in exploring materials informatics
optoelectronic semiconductors, specifically halide perovskites. Hu et al. [75] proposed and
in optoelectronic semiconductors, specifically halide perovskites. Hu et al. [75] proposed
developed [Link] (accessed on 19 August 2023), a web-based materials infor-
and developed [Link] (accessed on 19 August 2023), a web-based materials
matics toolbox. The MaterialsAtlas platform includes tools for chemical validity check,
informatics toolbox. The MaterialsAtlas platform includes tools for chemical validity check,
formation energy and e-above-hull energy check, property prediction, screening of hypo-
formation energy and e-above-hull energy check, property prediction, screening of hypo-
theticalmaterials,
thetical materials,and
andutility
utility tools.
tools. The
The toolbox
toolbox lowers
lowers thethe barrier
barrier for materials
for materials scientists
scientists in
in data-driven
data-driven exploratory
exploratory materials
materials discovery.
discovery.
[Link]
Figure Overviewofofthe
theJAMIP
JAMIPcode
codeframework.
[Link]
programcomprises
comprises three
three major
major parts
parts based
based
on the material data’s lifecycle: data generation (blue), data collection (yellow), and data learning
on the material data’s lifecycle: data generation (blue), data collection (yellow), and data learning
(green).Reproduced
(green). Reproducedwith
withpermission
permissionfrom
from[74].
[74].
4.
4. ML
ML ininMaterials
MaterialsScience
Science
4.1.
4.1. Prediction of MaterialProperties
Prediction of Material Properties
ML
MLhas
hasgained
gainedprominence
prominenceininrecent
recentyears
yearsininpredicting
predictingmaterial
material properties
propertiesdue to to
due
its advantages of high generalization ability and fast computational speed. It
its advantages of high generalization ability and fast computational speed. It has beenhas been
successfully applied to predict the structure, adsorption, electrical, catalytic, energy storage,
successfully applied to predict the structure, adsorption, electrical, catalytic, energy stor-
and thermodynamic properties of materials. The prediction results could even reach the
age, and thermodynamic properties of materials. The prediction results could even reach
same accuracy as high-fidelity models with low computational costs.
the same accuracy as high-fidelity models with low computational costs.
4.1.1. Molecular Properties
4.1.1. Molecular Properties
In the past, it was very time consuming to predict molecular properties based on high-
In the density
throughput past, it generalization
was very timecalculations.
consuming ML to predict molecular
allows fast properties
and accurate basedofon
prediction
high-throughput density generalization calculations. ML allows fast and
the structure or properties of molecules, compounds, and materials. In materials science, accurate predic-
tion of thefactors,
solubility structure
suchorasproperties
Hansen and of molecules,
Hildebrand compounds,
solubility, areandcritical
materials. In materials
parameters for
science, solubility
characterizing factors,properties
the physical such as Hansen andsubstances.
of various HildebrandKurotani
solubility, are
et al. critical
[76] parame-
successfully
ters for characterizing
developed the physical
a solubility prediction modelproperties of various
with a unique substances.
ML method, Kurotaniin-phase
the so-called et al. [76]
successfully
DNN (ip-DNN).developed a solubility
This algorithm prediction
started with themodel
analysiswith
of ainput
unique
dataML method,NMR
(including the so-
called in-phase
information, DNN index,
refractive (ip-DNN). This algorithm
and density). startedwas
The solubility withthen
thespeculated
analysis ofininput data
a multi-
step approach
(including NMR by information,
predicting intermediate elements,
refractive index, such as molecular
and density). components
The solubility was thenand spec-
molecular
ulated in adescriptors.
multi-stepAn intermediate
approach regression
by predicting model was also
intermediate utilized
elements, to improve
such the
as molecular
accuracy
components of the
andprediction.
molecularAdescriptors.
website dedicated to the established
An intermediate regressionsolubility
model was prediction
also uti-
methods has also the
lized to improve been developed,
accuracy of thewhich is available
prediction. free ofdedicated
A website charge. Liang
to the et al. [77]
established
proposed a generalized ML method based on ANNs to predict polymer compatibility
solubility prediction methods has also been developed, which is available free of charge. (the
total miscibility of polymers with each other at the molecular scale).
Liang et al. [77] proposed a generalized ML method based on ANNs to predict polymer The authors built a
database by collecting
compatibility (the totaldata from scattered
miscibility literature
of polymers withthrough natural
each other at thelanguage
molecular processing
scale). The
techniques. By using the proposed method, predictions could be made
authors built a database by collecting data from scattered literature through natural based on the basiclan-
guage processing techniques. By using the proposed method, predictions could be made
based on the basic molecular structure of the blended polymers and the blended
Materials 2023, 16, 5977 13 of 30
molecular structure of the blended polymers and the blended compositions (as an auxiliary).
This generalized approach yielded some results in illustrating polymer compatibility. A
prediction accuracy of no less than 75% was achieved on a dataset containing 1400 entries
in their model. Zeng et al. [78] developed an atomic table CNN that could predict the band
gap and ground energy. The model accuracy exceeded that of standard DFT calculations.
Furthermore, this model could accurately predict superconducting transition temperatures
and distinguish between superconductors and non-superconductors. With the help of
this model, 20 potential superconductor compounds with high superconducting transition
temperatures were screened out.
Figure 9.
Figure 9. Logic
Logic diagram
diagram of
of predicting
predicting the
the maximum
maximum energy
energy density
density and
and exploring
exploring the
the potential
potential
effective structure of composites through the ML method, reproduced with permission from [85].
effective structure of composites through the ML method, reproduced with permission from [85].
4.1.4.
4.1.4. Structural
Structural Health
Health
Structural
Structural health monitoring
health monitoring (SHM)
(SHM) utilizes
utilizes engineering,
engineering, scientific,
scientific, and
and foundational
foundational
knowledge
knowledge to prevent
prevent damage to property
property and
and life.
life. The core ofof the
the field
field of
of construction
construction
informatics
informatics isis the
the transmission,
transmission, processing,
processing, and
and visualization
visualization of architectural information,
of architectural information,
providing
providing effective
effective methods
methods forfor monitoring
monitoring structural
structural changes
changes [88,89]. ML provides
[88,89]. ML provides effec-
effec-
tive methods for monitoring structural changes. Dang et al. [90] proposed
tive methods for monitoring structural changes. Dang et al. [90] proposed a cloud-based a cloud-based
digital
digital twin
twin framework
framework for for SHM
SHM employing
employing deepdeep learning.
learning. The
The framework
framework consists
consists of
of
physical components, device measurements, and digital models formed
physical components, device measurements, and digital models formed by combining dif- by combining
different sub-models
ferent sub-models includingmathematical,
including mathematical,finite
finiteelement,
element,and
and MLML sub-models.
sub-models. The The data
data
interactions among the physical structure, digital model, and human interventions were
enhanced by using cloud computing infrastructure and a user-friendly web application.
Materials 2023, 16, 5977 15 of 30
interactions among the physical structure, digital model, and human interventions were
enhanced by using cloud computing infrastructure and a user-friendly web application.
The feasibility of the framework was demonstrated through case studies of the damage de-
tection of model bridges and real bridge structures utilizing deep learning algorithms, with
a high accuracy of 92%. Dong et al. [91] discussed the use of the eXtreme gradient boosting
(XGBoost) algorithm for predicting concrete electrical resistivity in SHM (Figure 10a). The
proposed XGBoost-algorithm-based prediction model considers all potential influencing
factors simultaneously. A database of 800 experimental instances was used to train and test
the model. The results showed that the XGBoost model achieved satisfactory predictive
performance. The study also identified the importance of curing age and cement content
in electrical resistivity measurement results. The XGBoost algorithm was chosen for its
high performance, ease of use, and better prediction accuracy than other algorithms. The
bond effect between the reinforcement and concrete guarantees the combined action of the
two materials. This is a critical factor that affects the mechanical properties of reinforced
concrete components and structures, e.g., bearing capacity and ductility [92]. Gao et al. [93]
developed a new solution for evaluating the bond strength of an FRP using AI-based
models. Two hybrid models, the imperialist competitive algorithm (ICA)-ANN and the
artificial bee colony (ABC)-ANN, were designed and compared. The results showed that
the ICA-ANN model had a higher predictive ability than the ABC-ANN model. The pro-
posed hybrid models can be used as a suitable substitute for empirical models in evaluating
FRP bond strength in concrete samples. Li et al. [94] utilized ML approaches to estimate
the bond strength between ultra-high-performance concrete (UHPC) and reinforcing bars.
A new database was created by integrating data from multiple published works. Nine
ML models, including linear models, tree models, and ANNs, were implemented to train
bond strength estimators based on the database. The results showed that the ANN and
RF models achieved the highest estimation performances, surpassing empirical formulas.
The study also analyzed the relative importance of different factors in determining bond
strength. Overall, the research provides a data-driven approach to estimating bond strength
and contributes to the understanding of bond performance between UHPC and reinforcing
bars. Su et al. [95] applied three ML approaches (multiple linear regression, SVM, and
ANN) to predict the interfacial bond strength between FRPs and concrete (Figure 10b). They
trained these models using two datasets containing experimental results from single-lap
shear tests, employed random search and grid search to find the optimal hyperparameters,
and analyzed input variables’ contributions using partial dependence plots. They also
developed a stacking strategy to improve prediction accuracy. The results showed that the
SVM approach had the best accuracy and efficiency. They concluded that ML methods
are feasible and efficient for predicting the bond strength of FRP laminates in reinforced
concrete structures.
most relevant for assessing nanomaterial toxicity and successfully correlated these prop-
erties with zebrafish physiological responses. It has been concluded that for the group of
metal and metal oxide nanomaterials, the core chemical composition, concentration, and
properties are influenced by the nanomaterial surface and medium composition (such as
zeta potential and agglomerate size), which have a significant impact on toxicity, even
though the ranking of different variables is subject to variation in the analytical method
and data model. Generalized nano-QSAR ensemble models offer a promising framework
for predicting the toxicity potential of new nanomaterials. Liu et al. [99] presented a meta-
analysis of phytosynthesized silver nanoparticles (AgNPs) with heterogeneous features
using DTs and RFs. The researchers found that exposure regime (including the time and
dose), plant family, and cell type were the most important predictors for cell viability for
green AgNPs. In addition, a discussion of the potential effects of major variables (cell
assays, inherent nanoparticle properties, and reaction parameters used in biosynthesis)
on AgNP-mediated cytotoxicity and model performance was presented to provide a basis
Materials 2023, 16, x FOR PEER REVIEW
for future research. The findings of this study may assist future studies in improving the 16 of 3
design of experiments and the development of virtual models or optimizations of green
AgNPs for specific applications.
Figure10.
Figure 10. (a)
(a) Schematic
SchematicofofXGBoost
XGBoost trees,
trees, adapted
adapted withwith permission
permission fromfrom
[91].[91]. (b) ML
(b) ML model con
model
struction process, adapted with permission from [95].
construction process, adapted with permission from [95].
Figure
Figure 11.11. Schematic
Schematic workflow
workflow of data compilation,
of data compilation, descriptor
descriptor generation, generation,
machine learningmachine
model- learn
ing, experimental
eling, validation,
experimental and mechanism
validation, interpretation,interpretation,
and mechanism reproduced with permission
reproduced fromwith
[97]. permiss
[97]. Adsorption Performance of Nanomaterials
4.1.6.
Because of their high surface area, ease of functionalization, and affinity toward a
4.1.6.
wide Adsorption
range Performance
of pollutants, nanomaterialsofare Nanomaterials
excellent adsorbents [100]. Moosavi et al. [101]
appliedBecause
four machine learning methods
of their high surface area, to model ease
dye adsorption on 16 activatedand
of functionalization, carbonaffinity t
adsorbents and determined the relationship between adsorption capacity and activated
wide range of pollutants, nanomaterials are excellent adsorbents [100]. Moosavi et
carbon parameters. The results indicated that agro-waste characteristics (pore volume,
applied
surface four
area, pH,machine
and particle learning methods
size) contributed to model
50.7% dye adsorption
to the adsorption [Link] 16 activated
Among
adsorbents
the agro-wasteand determined
characteristics, porethe relationship
volume and surface between
area were adsorption capacity and a
the most important
influencing variables, while
carbon parameters. Theparticle
resultssizeindicated
had a limited
thatimpact. With a hypothetical
agro-waste set of (pore
characteristics
approximately
surface area, pH, and particle size) contributed 50.7% to the adsorptionand
130,000 structures of metal–organic frameworks (MOFs) with methane efficiency
carbon dioxide adsorption data at different pressures, Guo et al. [102] established models
the agro-waste characteristics, pore volume and surface area were the most impo
for estimating gas adsorption capacities using two deep learning algorithms, multilayer
fluencing(MLPs)
perceptrons variables, while
and long particle
short-term size had
memory (LSTM) a limited
[Link].
The modelsWith
were a eval-
hypothetic
approximately
uated by performing130,000 structures
ten iterations of 10-foldofcross-validations
metal–organic andframeworks (MOFs) with
100 holdout validations.
The
and carbon dioxide adsorption data at different pressures, Guo etpredic-
performance of the MLP and LSTM models was similar with high accuracy of al. [102] est
tion. Those models that predicted gas adsorption at a higher pressure performed better
models for estimating gas adsorption capacities using two deep learning algorithm
than those that predicted gas adsorption at a lower pressure. In particular, deep learning
tilayerwere
models perceptrons
more accurate(MLPs) andmodels
than RF long reported
short-term memory
in the literature(LSTM) networks. The
when predicting
were
gas evaluated
adsorption by performing
capacities ten iterations
at low pressures. of 10-fold
Deep learning cross-validations
algorithms were found to beand 100
highly
validations. The performance of the MLP and LSTM models gas
effective in generating models capable of accurately predicting the wasadsorption
similar with hi
capacities of MOFs.
racy of prediction. Those models that predicted gas adsorption at a higher press
formed
4.2. better
Accelerated than those
Materials that
Synthesis andpredicted
Design gas adsorption at a lower pressure. In pa
deep learning models were more accurate than
In addition to being widely utilized for predicting RF models
material reported
properties, inplays
ML also the literatu
apredicting gas
pivotal role in theadsorption capacities
synthesis of new [Link]
low pressures. DeepML
the past few years, learning
has madealgorithm
Materials 2023, 16, 5977 18 of 30
significant progress in the exploration of novel materials, such as highly efficient molecular
organic light-emitting diodes [103], low thermal hysteresis shape memory alloys [104],
and piezoelectric materials with large electrical strain [105]. The use of ML for materials
synthesis not only significantly speeds up novel material discovery but also provides
insight into the basic composition changes in materials from big data.
Figure12.
Figure 12. Catalyst
Catalyststructures,
structures,target
target properties,
properties, andand computational
computational framework.
framework. (a) Structural rep
(a) Structural
resentation ofofthree-coordinated
representation three-coordinated andand four-coordinated
four-coordinated configurations.
configurations. Letter
Letter “M” “M” represents
represents the th
centralmetal
central metalatom,
atom,andand letter
letter “C”“C” represents
represents the coordinating
the coordinating atom ofatom ofTarget
M. (b) M. (b)properties
Target properties
for fo
describingthe
describing the 2 fixation
N2Nfixation performance
performance ofcatalyst.
of the the catalyst.
(c) ML(c) ML screening
screening and descriptor
and descriptor building buildin
frameworkofof
framework their
their work.
work. Reproduced
Reproduced withwith permission
permission from [112].
from [112].
Figure 13. (a) Workflow of the integrated model-based ML methods for accurate Tc prediction and
new superconductor material mining, adapted with permission from [117]. (b) A schematic layout of
the DeepSet architecture, adapted with permission from [119].
Figure 13. (a) Workflow of the integrated model-based ML methods for accurate Tc prediction and
new superconductor material mining, adapted with permission from [117]. (b) A schematic layout
Materials 2023, 16, 5977 of the DeepSet architecture, adapted with permission from [119]. 21 of 30
Figure14.
Figure [Link]
Schematicrepresentation
representationofofthe
theworking
workingflow
flowwhen
when machine
machine learning
learning models
models areare incor-
incorpo-
porated into the prediction of the crystallization propensity of MONCs, with permission from
rated into the prediction of the crystallization propensity of MONCs, with permission from [120]. [120].
process of identifying the most efficient recipe and reaction conditions is therefore time
consuming, laborious, and resource intensive [122]. In a recent study, Erick et al. [123] used
SVM classification and regression models to predict the synthesis of CsPbBr3 nanosheets
with controlled layer thicknesses. The SVM classification is shown to accurately predict the
likelihood that CsPbBr3 synthesis would form a majority population of quantum-confined
nanoplatelets. Additionally, SVM regression can be used to determine the average thickness
of the synthesis of CsPbBr3 nanoplatelets with sub-monolayer accuracy. Epps et al. [124]
proposed a method that is based on ML experiment selection and high-efficiency au-
tonomous flow chemistry. The approach utilized SVM regression to predict the thickness of
the nanoplatelets and was shown to be accurate and reliable. Using this method, inorganic
perovskite quantum dots (QDs) in flow were synthesized autonomously. By using less than
210 mL of starting solutions and without user selection, this method synthesized precision
tailored QD compositions within 30 h. This would enable the commercialization of these
QDs, as well as their integration into various applications. Furthermore, the method could
be used for other types of nanomaterials, such as nanorods and nanowires.
Figure 15.
Figure 15. Structures
Structures ofof machine
machine learning
learning models
models for
for predicting
predicting optical
optical properties
properties and
and designing
designing
nanoparticles. (a) Far- and near-field optical data obtained from the finite-difference time-domain
nanoparticles. (a) Far- and near-field optical data obtained from the finite-difference time-domain
(FDTD) simulations were used to train three different machine learning models: far-field spectra
(FDTD) simulations were used to train three different machine learning models: far-field spectra
and structural information for (i) structure classification, far-field spectra and dimensions for (ii) the
and structural information for (i) structure classification, far-field spectra and dimensions for (ii) the
spectral DNN, and near-field enhancement maps and dimensions for (iii) the E-field DNN. After
spectral
training,DNN,
machine and near-field
learning enhancement
models can be usedmaps and dimensions
to perform for (iii) the
forward prediction E-field
and/or [Link].
inverse After
training, machine learning models can be used to perform forward prediction and/or inverse
The solid and dashed red arrows represent the forward prediction and the inverse design process, design.
The solid and(b)
respectively. dashed red architecture
Detailed arrows represent
of thethe forward
three machineprediction
learningand the inverse
models design
in Figure process,
9a, with per-
mission from [128].
respectively. (b) Detailed architecture of the three machine learning models in Figure 9a, with
permission from [128].
5. Conclusions, Challenges, and Prospects
5. Conclusions, Challenges, and Prospects
This review discussed the use of machine learning (ML) in the field of materials sci-
ence This review discussed
for predicting materialthe use of machine
properties learning
and guiding (ML) insynthesis.
material the field ofThe
materials
reviewscience
briefly
for predicting material properties and guiding material synthesis. The
outlined the basic principles of ML and introduced commonly used algorithms and review briefly
their
outlined the basic
applications principles
in material of MLand
screening andproperty
introduced commonly
prediction. used
It also algorithms
presented theand their
research
applications in material screening and property prediction. It also presented the research
progress of ML in predicting material properties and guiding material synthesis. The
Materials 2023, 16, 5977 24 of 30
progress of ML in predicting material properties and guiding material synthesis. The review
suggested that ML can greatly reduce computational costs, shorten the development cycle,
and improve computational accuracy, making it a promising research approach in novel
materials screening and material property prediction.
It is important to note, however, that the following challenges still exist. Most ML
algorithms require large amounts of data to work properly. Even for the simplest problems,
thousands of examples are desired. Acquiring an effective dataset is critical for the research
and implementation of ML in materials science. However, data in materials science are
characterized by high acquisition costs, excessive concentration or dispersion, and a lack
of uniform processing standards. A dataset with a large amount of data, a uniform distri-
bution, and matching feature parameters is often extremely difficult to obtain. Although
material databases have greatly facilitated researchers’ access to data, many published data
have not been specified to date. The task of enriching existing databases is challenging. Text
mining techniques could be effective in rapidly collecting data scattered in the literature.
This approach could greatly enhance existing databases and create specialized databases.
The selection of features significantly affects the accuracy of ML models. Currently, the
use of manual feature engineering to filter features is often influenced by the researcher’s
experience and intuition. This approach may overlook some significant features. In contrast,
automated feature engineering automatically constructs new candidate features from the
data and selects the most appropriate features for model training, which could effectively
solve the current dilemma.
ML methods cannot replace traditional computational and experimental studies. Al-
though ML methods have shown remarkable promise in guiding the synthesis of novel
materials and predicting material properties, they are still mostly “black boxes” [108]. The
predicted results still need to be experimentally verified and the underlying physicochemi-
cal laws still need to be studied in depth. Therefore, ML can only perform some exploratory
tasks at present. With further improvement of theories and methods, however, ML might
eventually replace traditional experimental research by providing novel ideas and research
methods for the field of materials science. The application of ML in the field of materials
science and engineering is just the beginning, and its potential is endless in the future.
Author Contributions: Writing—original draft preparation, G.H. and Y.G.; writing—review and
editing, Z.N. and Y.C. All authors have read and agreed to the published version of the manuscript.
Funding: This work was supported by the Postgraduate Research and Practice Innovation Program of
Jiangsu Province (College Project), China, the Natural Science Foundation of Jiangsu Province, China
(Grant No. BK20200686), and the National Natural Science Foundation of China (Grant No. 52206257).
Conflicts of Interest: The authors declare that the research was conducted in the absence of any
commercial or financial relationships that could be construed as potential conflicts of interest.
Abbreviations
ABC Artificial bee colony
AFLOW Automatic Flow
AgNP Silver nanoparticle
AI Artificial intelligence
ANN Artificial neural network
CNN Convolutional neural network
COD Crystallography Open Database
CSD Cambridge Structural Database
DBM Deep Boltzmann machine
DBN Deep belief network
DFT Density functional theory
DNN Deep neural network
Materials 2023, 16, 5977 25 of 30
DT Decision tree
FDTD Finite-difference time-domain
FRP Fiber-reinforced polymer
GAN Generative adversarial network
GBRT Gradient boosted regression tree
GGA Generalized gradient approximation
GLOnet Global optimization network
HEA High-entropy alloy
HHV Higher heating value
ICA Imperialist competitive algorithm
ICSD Inorganic Crystal Structure Database
JAMIP Jilin Artificial-intelligence aided Materials-design Integrated Package
KNN K-nearest neighbor
LSTM Long short-term memory
MDN Mixture density network
ML Machine learning
MLP Multilayer perceptron
MOF Metal–organic framework
MONC Metal–organic nanocapsule
NMR Nuclear magnetic resonance
OMDB Organic Materials Database
OQMD Open Quantum Materials Database
QD Quantum dot
QSAR Quantitative structure–activity relationship
R-CNN Region-based CNN
ResNet Residual network
RF Random forest
RMSE Root mean square error
RNN Recurrent neural network
SEM Scanning electron microscope
SHM Structural health monitoring
SMILE Simplified molecular input line entry
SVM Support vector machine
SVR Support vector regression
TEM Transmission electron microscope
TGNN Tupleswise graph neural network
UHPC Ultra-high-performance concrete
XGBoost eXtreme gradient boosting
XPS X-ray photoelectron spectroscopy
XRD X-ray diffraction
µ-PTAMAM Micro-plasma transfer arc metal additive manufacturing
References
1. Lu, S.; Zhou, Q.; Ouyang, Y.; Guo, Y.; Li, Q.; Wang, J. Accelerated discovery of stable lead-free hybrid organic-inorganic
perovskites via machine learning. Nat. Commun. 2018, 9, 3405. [CrossRef] [PubMed]
2. Kolahalam, L.A.; Viswanath, I.K.; Diwakar, B.S.; Govindh, B.; Reddy, V.; Murthy, Y. Review on nanomaterials: Synthesis and
applications. Mater. Today Proc. 2019, 18, 2182–2190. [CrossRef]
3. Schleder, G.R.; Padilha, A.C.; Acosta, C.M.; Costa, M.; Fazzio, A. From DFT to machine learning: Recent approaches to materials
science–a review. J. Phys. Mater. 2019, 2, 032001. [CrossRef]
4. Butler, K.T.; Davies, D.W.; Cartwright, H.; Isayev, O.; Walsh, A. Machine learning for molecular and materials science. Nature
2018, 559, 547–555. [CrossRef]
5. Chibani, S.; Coudert, F.-X. Machine learning approaches for the prediction of materials properties. APL Mater. 2020, 8, 080701.
[CrossRef]
6. Rajendra, P.; Girisha, A.; Naidu, T.G. Advancement of machine learning in materials science. Mater. Today Proc. 2022, 62, 5503–5507.
[CrossRef]
7. Ruoff, R.; Tse, D.S.; Malhotra, R.; Lorents, D.C. Solubility of fullerene (C60) in a variety of solvents. J. Phys. Chem. 1993, 97, 3379–3383.
[CrossRef]
Materials 2023, 16, 5977 26 of 30
8. Guo, K.; Yang, Z.; Yu, C.-H.; Buehler, M.J. Artificial intelligence and machine learning in design of mechanical materials.
Mater. Horiz. 2021, 8, 1153–1172. [CrossRef]
9. Cai, J.; Chu, X.; Xu, K.; Li, H.; Wei, J. Machine learning-driven new material discovery. Nanoscale Adv. 2020, 2, 3115–3130.
[CrossRef]
10. Fang, J.; Xie, M.; He, X.; Zhang, J.; Hu, J.; Chen, Y.; Yang, Y.; Jin, Q. Machine learning accelerates the materials discovery.
Mater. Today Commun. 2022, 33, 104900. [CrossRef]
11. Chen, A.; Zhang, X.; Zhou, Z. Machine learning:Accelerating materials development for energy storage and conversion. InfoMat
2020, 2, 553–576. [CrossRef]
12. Liu, Y.; Niu, C.; Wang, Z.; Gan, Y.; Zhu, Y.; Sun, S.; Shen, T. Machine learning in materials genome initiative: A review. J. Mater.
Sci. Technol. 2020, 57, 113–122. [CrossRef]
13. Raccuglia, P.; Elbert, K.C.; Adler, P.D.; Falk, C.; Wenny, M.B.; Mollo, A.; Zeller, M.; Friedler, S.A.; Schrier, J.; Norquist, A.J.
Machine-learning-assisted materials discovery using failed experiments. Nature 2016, 533, 73–76. [CrossRef] [PubMed]
14. Zhou, L.; Yao, A.M.; Wu, Y.; Hu, Z.; Huang, Y.; Hong, Z. Machine Learning Assisted Prediction of Cathode Materials for Zn-Ion
Batteries. Adv. Theory Simul. 2021, 4, 2100196. [CrossRef]
15. Ridzuan, F.; Zainon, W.M.N.W. A review on data cleansing methods for big data. Procedia Comput. Sci. 2019, 161, 731–738.
[CrossRef]
16. Hossen, M.S. Data preprocess. Machine Learning and Big Data: Concepts, Algorithms, Tools and Applications; Scrivener Publishing:
Beverly, MA, USA, 2020; pp. 71–103.
17. Wu, Y.-W.; Tang, Y.-H.; Tringe, S.G.; Simmons, B.A.; Singer, S.W. MaxBin: An automated binning method to recover individual
genomes from metagenomes using an expectation-maximization algorithm. Microbiome 2014, 2, 26. [CrossRef]
18. Fernández-Delgado, M.; Sirsat, M.S.; Cernadas, E.; Alawadi, S.; Barro, S.; Febrero-Bande, M. An extensive experimental survey of
regression methods. Neural Netw. 2019, 111, 11–34. [CrossRef]
19. Liu, G.-H.; Shen, H.-B.; Yu, D.-J. Prediction of protein–protein interaction sites with machine-learning-based data-cleaning and
post-filtering procedures. J. Membr. Biol. 2016, 249, 141–153. [CrossRef]
20. Wei, J.; Chu, X.; Sun, X.Y.; Xu, K.; Deng, H.X.; Chen, J.; Wei, Z.; Lei, M. Machine learning in materials science. InfoMat 2019,
1, 338–358. [CrossRef]
21. Wang, M.; Wang, T.; Cai, P.; Chen, X. Nanomaterials Discovery and Design through Machine Learning. Small Methods 2019,
3, 1900025. [CrossRef]
22. Schmidt, J.; Marques, M.R.G.; Botti, S.; Marques, M.A.L. Recent advances and applications of machine learning in solid-state
materials science. NPJ Comput. Mater. 2019, 36, 83. [CrossRef]
23. Hou, Y.; Wang, Q.; Tan, T. Prediction of carbon dioxide emissions in China using shallow learning with cross validation. Energies
2022, 15, 8642. [CrossRef]
24. Kurani, A.; Doshi, P.; Vakharia, A.; Shah, M. A comprehensive comparative study of artificial neural network (ANN) and support
vector machines (SVM) on stock forecasting. Ann. Data Sci. 2023, 10, 183–208. [CrossRef]
25. Cover, T.M. Rates of convergence for nearest neighbor procedures. In Proceedings of the Hawaii International Conference on
Systems Sciences, Honolulu, HI, USA, 29–30 January 1968.
26. Peterson, L.E. K-nearest neighbor. Scholarpedia 2009, 4, 1883. [CrossRef]
27. Sharma, A.; Madhushri, P.; Kushvaha, V. Dynamic fracture toughness prediction of fiber/epoxy composites using K-nearest
neighbor (KNN) method. In Handbook of Epoxy/Fiber Composites; Springer: Berlin/Heidelberg, Germany, 2022; pp. 1–16.
28. Sun, B.; Du, J.; Gao, T. Study on the improvement of K-nearest-neighbor algorithm. In Proceedings of the 2009 International
Conference on Artificial Intelligence and Computational Intelligence, Shanghai, China, 7–8 November 2009; pp. 390–393.
29. Hunt, E. Concept Learning: An Information Processing Problem; John Wiley & Sons, Inc.: Hoboken, NJ, USA, 1962.
30. Mak, B.; Munakata, T. Rule extraction from expert heuristics: A comparative study of rough sets with neural networks and ID3.
Eur. J. Oper. Res. 2002, 136, 212–229. [CrossRef]
31. Ruggieri, S. Efficient C4. 5 [classification algorithm]. IEEE Trans. Knowl. Data Eng. 2002, 14, 438–444. [CrossRef]
32. Rokach, L.; Maimon, O. Top-down induction of decision trees classifiers-a survey. IEEE Trans. Syst. Man Cybern. Part C (Appl.
Rev.) 2005, 35, 476–487. [CrossRef]
33. Liu, X.; Liu, T.; Feng, P. Long-term performance prediction framework based on XGBoost decision tree for pultruded FRP
composites exposed to water, humidity and alkaline solution. Compos. Struct. 2022, 284, 115184. [CrossRef]
34. Liu, Y.; Wang, Y.; Zhang, J. New machine learning algorithm: Random forest. In Proceedings of the Information Computing and
Applications: Third International Conference, ICICA 2012, Chengde, China, 14–16 September 2012; pp. 246–252.
35. McCulloch, W.S.; Pitts, W. A logical calculus of the ideas immanent in nervous activity. Bull. Math. Biophys. 1943, 5, 115–133.
[CrossRef]
36. Wu, Y.-C.; Feng, J.-W. Development and application of artificial neural network. Wirel. Pers. Commun. 2018, 102, 1645–1656.
[CrossRef]
37. Huang, Y. Advances in artificial neural networks–methodological development and application. Algorithms 2009, 2, 973–1007.
[CrossRef]
38. Abiodun, O.I.; Jantan, A.; Omolara, A.E.; Dada, K.V.; Mohamed, N.A.; Arshad, H. State-of-the-art in artificial neural network
applications: A survey. Heliyon 2018, 4, e00938. [CrossRef]
Materials 2023, 16, 5977 27 of 30
39. Hmede, R.; Chapelle, F.; Lapusta, Y. Review of neural network modeling of shape memory alloys. Sensors 2022, 22, 5610.
[CrossRef]
40. Mendizabal, A.; Márquez-Neila, P.; Cotin, S. Simulation of hyperelastic materials in real-time using deep learning. Med. Image
Anal. 2020, 59, 101569. [CrossRef] [PubMed]
41. Savaedi, Z.; Motallebi, R.; Mirzadeh, H. A review of hot deformation behavior and constitutive models to predict flow stress of
high-entropy alloys. J. Alloys Compd. 2022, 903, 163964. [CrossRef]
42. Sharma, A.; Madhushri, P.; Kushvaha, V.; Kumar, A. Prediction of the fracture toughness of silicafilled epoxy composites
using K-nearest neighbor (KNN) method. In Proceedings of the 2020 International Conference on Computational Performance
Evaluation (ComPE), Shillong, India, 2–4 July 2020; pp. 194–198.
43. Kumar, P.; Jain, N.K. Surface roughness prediction in micro-plasma transferred arc metal additive manufacturing process using
K-nearest neighbors algorithm. Int. J. Adv. Manuf. Technol. 2022, 119, 2985–2997. [CrossRef]
44. Ghouchan Nezhad Noor Nia, R.; Jalali, M.; Houshmand, M. A Graph-Based k-Nearest Neighbor (KNN) Approach for Predicting
Phases in High-Entropy Alloys. Appl. Sci. 2022, 12, 8021. [CrossRef]
45. Wang, R.; Zheng, Z.; Yin, Z.; Wang, Y. Identification Method of Transformer Winding Material Based on Support Vector Machine.
In Proceedings of the 2022 2nd International Conference on Electrical Engineering and Control Science (IC2ECS), Nanjing, China,
16–18 December 2022; pp. 913–917.
46. Martinez, R.F.; Jimbert, P.; Callejo, L.M.; Barbero, J.I. Material Fracture Life Prediction Under High Temperature Creep Conditions
Using Support Vector Machines And Artificial Neural Networks Techniques. In Proceedings of the 2021 31st International
Conference on Computer Theory and Applications (ICCTA), Alexandria, Egypt, 11–13 December 2021; pp. 127–132.
47. Ahmad, M.; Hu, J.-L.; Ahmad, F.; Tang, X.-W.; Amjad, M.; Iqbal, M.J.; Asim, M.; Farooq, A. Supervised learning methods for
modeling concrete compressive strength prediction at high temperature. Materials 2021, 14, 1983. [CrossRef]
48. Sun, Y.; Li, G.; Zhang, N.; Chang, Q.; Xu, J.; Zhang, J. Development of ensemble learning models to evaluate the strength of
coal-grout materials. Int. J. Min. Sci. Technol. 2021, 31, 153–162. [CrossRef]
49. Samadi, S.H.; Ghobadian, B.; Nosrati, M. Prediction of higher heating value of biomass materials based on proximate analysis
using gradient boosted regression trees method. Energy Sources Part A Recovery Util. Environ. Eff. 2021, 43, 672–681. [CrossRef]
50. Shahmansouri, A.A.; Yazdani, M.; Ghanbari, S.; Bengar, H.A.; Jafari, A.; Ghatte, H.F. Artificial neural network model to predict
the compressive strength of eco-friendly geopolymer concrete incorporating silica fume and natural zeolite. J. Clean. Prod. 2021,
279, 123697. [CrossRef]
51. Liu, Q.-F.; Iqbal, M.F.; Yang, J.; Lu, X.-Y.; Zhang, P.; Rauf, M. Prediction of chloride diffusivity in concrete using artificial neural
network: Modelling and performance evaluation. Constr. Build. Mater. 2021, 268, 121082. [CrossRef]
52. Hinton, G.E.; Salakhutdinov, R.R. Reducing the dimensionality of data with neural networks. Science 2006, 313, 504–507.
[CrossRef] [PubMed]
53. Du, X.; Cai, Y.; Wang, S.; Zhang, L. Overview of deep learning. In Proceedings of the 2016 31st Youth Academic Annual
Conference of Chinese Association of Automation (YAC), Wuhan, China, 11–13 November 2016; pp. 159–164.
54. Agrawal, A.; Choudhary, A. Deep materials informatics: Applications of deep learning in materials science. MRS Commun. 2019,
9, 779–792. [CrossRef]
55. Gu, F.; Khoshelham, K.; Yu, C.; Shang, J. Accurate step length estimation for pedestrian dead reckoning localization using stacked
autoencoders. IEEE Trans. Instrum. Meas. 2018, 68, 2705–2713. [CrossRef]
56. Yang, H.; Shen, S.; Yao, X.; Sheng, M.; Wang, C. Competitive deep-belief networks for underwater acoustic target recognition.
Sensors 2018, 18, 952. [CrossRef]
57. Duong, C.N.; Luu, K.; Quach, K.G.; Bui, T.D. Deep appearance models: A deep boltzmann machine approach for face modeling.
Int. J. Comput. Vis. 2019, 127, 437–455. [CrossRef]
58. Parashar, A.; Raina, P.; Shao, Y.S.; Chen, Y.-H.; Ying, V.A.; Mukkara, A.; Venkatesan, R.; Khailany, B.; Keckler, S.W.; Emer, J.
Timeloop: A systematic approach to dnn accelerator evaluation. In Proceedings of the 2019 IEEE International Symposium on
Performance Analysis of Systems and Software (ISPASS), Madison, WI, USA, 24–26 March 2019; pp. 304–315.
59. Gu, J.; Wang, Z.; Kuen, J.; Ma, L.; Shahroudy, A.; Shuai, B.; Liu, T.; Wang, X.; Wang, G.; Cai, J. Recent advances in convolutional
neural networks. Pattern Recognit. 2018, 77, 354–377. [CrossRef]
60. LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [CrossRef]
61. Wu, S.-w.; Yang, J.; Cao, G.-m. Prediction of the Charpy V-notch impact energy of low carbon steel using a shallow neural network
and deep learning. Int. J. Miner. Metall. Mater. 2021, 28, 1309–1320. [CrossRef]
62. Sun, W.; Li, M.; Li, Y.; Wu, Z.; Sun, Y.; Lu, S.; Xiao, Z.; Zhao, B.; Sun, K. The use of deep learning to fast evaluate organic
photovoltaic materials. Adv. Theory Simul. 2019, 2, 1800116. [CrossRef]
63. Konno, T.; Kurokawa, H.; Nabeshima, F.; Sakishita, Y.; Ogawa, R.; Hosako, I.; Maeda, A. Deep learning model for finding new
superconductors. Phys. Rev. B 2021, 103, 014509. [CrossRef]
64. Li, C.; Wang, C.; Sun, M.; Zeng, Y.; Yuan, Y.; Gou, Q.; Wang, G.; Guo, Y.; Pu, X. Correlated RNN Framework to Quickly Generate
Molecules with Desired Properties for Energetic Materials in the Low Data Regime. J. Chem. Inf. Model. 2022, 62, 4873–4887.
[CrossRef] [PubMed]
65. Zhang, H.; Wang, Y.; Zhao, H.; Lu, K.; Yu, D.; Wen, J. Accelerated topological design of metaporous materials of broadband sound
absorption performance by generative adversarial networks. Mater. Des. 2021, 207, 109855. [CrossRef]
Materials 2023, 16, 5977 28 of 30
66. Unni, R.; Yao, K.; Zheng, Y. Deep convolutional mixture density network for inverse design of layered photonic structures.
ACS Photonics 2020, 7, 2703–2712. [CrossRef]
67. Huang, K.; Lei, H.; Jiao, Z.; Zhong, Z. Recycling waste classification using vision transformer on portable device. Sustainability
2021, 13, 11572. [CrossRef]
68. Jiang, J.; Fan, J.A. Multiobjective and categorical global optimization of photonic structures based on ResNet generative neural
networks. Nanophotonics 2020, 10, 361–369. [CrossRef]
69. Wang, M.; Yang, L.; Zhao, Z.; Guo, Y. Intelligent prediction of wear location and mechanism using image identification based on
improved Faster R-CNN model. Tribol. Int. 2022, 169, 107466. [CrossRef]
70. Ramprasad, R.; Batra, R.; Pilania, G.; Mannodi-Kanakkithodi, A.; Kim, C. Machine learning in materials informatics: Recent
applications and prospects. NPJ Comput. Mater. 2017, 3, 54. [CrossRef]
71. Li, M.; Zhang, H.; Li, S.; Zhu, W.; Ke, Y. Machine learning and materials informatics approaches for predicting transverse
mechanical properties of unidirectional CFRP composites with microvoids. Mater. Des. 2022, 224, 111340. [CrossRef]
72. Ramakrishna, S.; Zhang, T.-Y.; Lu, W.-C.; Qian, Q.; Low, J.S.C.; Yune, J.H.R.; Tan, D.Z.L.; Bressan, S.; Sanvito, S.; Kalidindi, S.R.
Materials informatics. J. Intell. Manuf. 2019, 30, 2307–2326. [CrossRef]
73. Al-Saban, O.; Abdellatif, S.O. Optoelectronic materials informatics: Utilizing random-forest machine learning in optimizing
the harvesting capabilities of mesostructured-based solar cells. In Proceedings of the 2021 International Telecommunications
Conference (ITC-Egypt), Alexandria, Egypt, 13–15 July 2021; pp. 1–4.
74. Zhao, X.-G.; Zhou, K.; Xing, B.; Zhao, R.; Luo, S.; Li, T.; Sun, Y.; Na, G.; Xie, J.; Yang, X. JAMIP: An artificial-intelligence aided
data-driven infrastructure for computational materials informatics. Sci. Bull. 2021, 66, 1973–1985. [CrossRef] [PubMed]
75. Hu, J.; Stefanov, S.; Song, Y.; Omee, S.S.; Louis, S.-Y.; Siriwardane, E.M.; Zhao, Y.; Wei, L. MaterialsAtlas. org: A materials
informatics web app platform for materials discovery and survey of state-of-the-art. NPJ Comput. Mater. 2022, 8, 65. [CrossRef]
76. Kurotani, A.; Kakiuchi, T.; Kikuchi, J. Solubility prediction from molecular properties and analytical data using an in-phase deep
neural network (Ip-DNN). ACS Omega 2021, 6, 14278–14287. [CrossRef]
77. Liang, Z.; Li, Z.; Zhou, S.; Sun, Y.; Yuan, J.; Zhang, C. Machine-learning exploration of polymer compatibility. Cell Rep. Phys. Sci.
2022, 3, 100931. [CrossRef]
78. Zeng, S.; Zhao, Y.; Li, G.; Wang, R.; Wang, X.; Ni, J. Atom table convolutional neural networks for an accurate prediction of
compounds properties. NPJ Comput. Mater. 2019, 5, 84. [CrossRef]
79. Venkatraman, V. The utility of composition-based machine learning models for band gap prediction. Comput. Mater. Sci. 2021,
197, 110637. [CrossRef]
80. Xu, P.; Lu, T.; Ju, L.; Tian, L.; Li, M.; Lu, W. Machine Learning Aided Design of Polymer with Targeted Band Gap Based on DFT
Computation. J. Phys. Chem. B 2021, 125, 601–611. [CrossRef]
81. Espinosa, R.; Ponce, H.; Ortiz-Medina, J. A 3D orthogonal vision-based band-gap prediction using deep learning: A proof of
concept. Comput. Mater. Sci. 2022, 202, 110967. [CrossRef]
82. Wang, T.; Zhang, K.; Thé, J.; Yu, H. Accurate prediction of band gap of materials using stacking machine learning model.
Comput. Mater. Sci. 2022, 201, 110899. [CrossRef]
83. Na, G.S.; Jang, S.; Lee, Y.-L.; Chang, H. Tuplewise material representation based machine learning for accurate band gap prediction.
J. Phys. Chem. A 2020, 124, 10616–10623. [CrossRef] [PubMed]
84. Shen, Z.H.; Liu, H.X.; Shen, Y.; Hu, J.M.; Chen, L.Q.; Nan, C.W. Machine learning in energy storage materials. Interdiscip. Mater.
2022, 1, 175–195. [CrossRef]
85. Feng, Y.; Tang, W.; Zhang, Y.; Zhang, T.; Shang, Y.; Chi, Q.; Chen, Q.; Lei, Q. Machine learning and microstructure design of
polymer nanocomposites for energy storage application. High Volt. 2022, 7, 242–250. [CrossRef]
86. Yue, D.; Feng, Y.; Liu, X.X.; Yin, J.H.; Zhang, W.C.; Guo, H.; Su, B.; Lei, Q.Q. Prediction of Energy Storage Performance in Polymer
Composites Using High-Throughput Stochastic Breakdown Simulation and Machine Learning. Adv. Sci. 2022, 9, 2105773.
[CrossRef] [PubMed]
87. Ojih, J.; Onyekpe, U.; Rodriguez, A.; Hu, J.; Peng, C.; Hu, M. Machine Learning Accelerated Discovery of Promising Thermal
Energy Storage Materials with High Heat Capacity. ACS Appl. Mater. Interfaces 2022, 14, 43277–43289. [CrossRef] [PubMed]
88. Malekloo, A.; Ozer, E.; AlHamaydeh, M.; Girolami, M. Machine learning and structural health monitoring overview with
emerging technology and high-dimensional data source highlights. Struct. Health Monit. 2022, 21, 1906–1955. [CrossRef]
89. Cao, Y.; Miraba, S.; Rafiei, S.; Ghabussi, A.; Bokaei, F.; Baharom, S.; Haramipour, P.; Assilzadeh, H. Economic application of
structural health monitoring and internet of things in efficiency of building information modeling. Smart Struct. Syst. 2020,
26, 559–573.
90. Dang, H.V.; Tatipamula, M.; Nguyen, H.X. Cloud-based digital twinning for structural health monitoring using deep learning.
IEEE Trans. Ind. Inform. 2021, 18, 3820–3830. [CrossRef]
91. Dong, W.; Huang, Y.; Lehane, B.; Ma, G. XGBoost algorithm-based prediction of concrete electrical resistivity for structural health
monitoring. Autom. Constr. 2020, 114, 103155. [CrossRef]
92. Fu, B.; Chen, S.-Z.; Liu, X.-R.; Feng, D.-C. A probabilistic bond strength model for corroded reinforced concrete based on weighted
averaging of non-fine-tuned machine learning models. Constr. Build. Mater. 2022, 318, 125767. [CrossRef]
93. Gao, J.; Koopialipoor, M.; Armaghani, D.J.; Ghabussi, A.; Baharom, S.; Morasaei, A.; Shariati, A.; Khorami, M.; Zhou, J. Evaluating
the bond strength of FRP in concrete samples using machine learning methods. Smart Struct. Syst. Int. J. 2020, 26, 403–418.
Materials 2023, 16, 5977 29 of 30
94. Li, Z.; Qi, J.; Hu, Y.; Wang, J. Estimation of bond strength between UHPC and reinforcing bars using machine learning approaches.
Eng. Struct. 2022, 262, 114311. [CrossRef]
95. Su, M.; Zhong, Q.; Peng, H.; Li, S. Selected machine learning approaches for predicting the interfacial bond strength between
FRPs and concrete. Constr. Build. Mater. 2021, 270, 121456. [CrossRef]
96. Khan, B.M.; Cohen, Y. Predictive Nanotoxicology: Nanoinformatics Approach to Toxicity Analysis of Nanomaterials. In Machine
Learning in Chemical Safety and Health: Fundamentals with Applications; John Wiley & Sons: Hoboken, NJ, USA, 2022; pp. 199–250.
97. Huang, Y.; Li, X.; Cao, J.; Wei, X.; Li, Y.; Wang, Z.; Cai, X.; Li, R.; Chen, J. Use of dissociation degree in lysosomes to predict metal
oxide nanoparticle toxicity in immune cells: Machine learning boosts nano-safety assessment. Environ. Int. 2022, 164, 107258.
[CrossRef] [PubMed]
98. Gousiadou, C.; Marchese Robinson, R.; Kotzabasaki, M.; Doganis, P.; Wilkins, T.; Jia, X.; Sarimveis, H.; Harper, S. Machine learning
predictions of concentration-specific aggregate hazard scores of inorganic nanomaterials in embryonic zebrafish. Nanotoxicology
2021, 15, 446–476. [CrossRef] [PubMed]
99. Liu, L.; Zhang, Z.; Cao, L.; Xiong, Z.; Tang, Y.; Pan, Y. Cytotoxicity of phytosynthesized silver nanoparticles: A meta-analysis by
machine learning algorithms. Sustain. Chem. Pharm. 2021, 21, 100425. [CrossRef]
100. Sajid, M.; Ihsanullah, I.; Khan, M.T.; Baig, N. Nanomaterials-based adsorbents for remediation of microplastics and nanoplastics
in aqueous media: A review. Sep. Purif. Technol. 2022, 305, 122453. [CrossRef]
101. Moosavi, S.; Manta, O.; El-Badry, Y.A.; Hussein, E.E.; El-Bahy, Z.M.; Mohd Fawzi, N.f.B.; Urbonavičius, J.; Moosavi, S.M.H.
A study on machine learning methods’ application for dye adsorption prediction onto agricultural waste activated carbon.
Nanomaterials 2021, 11, 2734. [CrossRef]
102. Guo, W.; Liu, J.; Dong, F.; Chen, R.; Das, J.; Ge, W.; Xu, X.; Hong, H. Deep learning models for predicting gas adsorption capacity
of nanomaterials. Nanomaterials 2022, 12, 3376. [CrossRef]
103. Gómez-Bombarelli, R.; Aguilera-Iparraguirre, J.; Hirzel, T.D.; Duvenaud, D.; Maclaurin, D.; Blood-Forsythe, M.A.; Chae, H.S.;
Einzinger, M.; Ha, D.-G.; Wu, T. Design of efficient molecular organic light-emitting diodes by a high-throughput virtual screening
and experimental approach. Nat. Mater. 2016, 15, 1120–1127. [CrossRef]
104. Xue, D.; Yuan, R.; Zhou, Y.; Xue, D.; Lookman, T.; Zhang, G.; Ding, X.; Sun, J. Design of high temperature Ti-Pd-Cr shape memory
alloys with small thermal hysteresis. Sci. Rep. 2016, 6, 28244. [CrossRef] [PubMed]
105. Li, W.; Yang, T.; Liu, C.; Huang, Y.; Chen, C.; Pan, H.; Xie, G.; Tai, H.; Jiang, Y.; Wu, Y. Optimizing piezoelectric nanocomposites by
high-throughput phase-field simulation and machine learning. Adv. Sci. 2022, 9, 2105550. [CrossRef] [PubMed]
106. Zhang, L.; He, M.; Shao, S. Machine learning for halide perovskite materials. Nano Energy 2020, 78, 105380. [CrossRef]
107. Li, L.; Tao, Q.; Xu, P.; Yang, X.; Lu, W.; Li, M. Studies on the regularity of perovskite formation via machine learning. Comput. Mater.
Sci. 2021, 199, 110712. [CrossRef]
108. Liu, H.; Cheng, J.; Dong, H.; Feng, J.; Pang, B.; Tian, Z.; Ma, S.; Xia, F.; Zhang, C.; Dong, L. Screening stable and metastable ABO3
perovskites using machine learning and the materials project. Comput. Mater. Sci. 2020, 177, 109614. [CrossRef]
109. Omprakash, P.; Manikandan, B.; Sandeep, A.; Shrivastava, R.; Viswesh, P.; Panemangalore, D.B. Graph representational learning
for bandgap prediction in varied perovskite crystals. Comput. Mater. Sci. 2021, 196, 110530. [CrossRef]
110. Wang, Z.; Cai, J.; Wang, Q.; Wu, S.; Li, J. Unsupervised discovery of thin-film photovoltaic materials from unlabeled data.
NPJ Comput. Mater. 2021, 7, 128. [CrossRef]
111. Huang, K.; Zhan, X.-L.; Chen, F.-Q.; Lü, D.-W. Catalyst design for methane oxidative coupling by using artificial neural network
and hybrid genetic algorithm. Chem. Eng. Sci. 2003, 58, 81–87. [CrossRef]
112. Zhang, S.; Lu, S.; Zhang, P.; Tian, J.; Shi, L.; Ling, C.; Zhou, Q.; Wang, J. Accelerated Discovery of Single-Atom Catalysts for
Nitrogen Fixation via Machine Learning. Energy Environ. Mater. 2023, 6, e12304. [CrossRef]
113. Wei, S.; Baek, S.; Yue, H.; Liu, M.; Yun, S.J.; Park, S.; Lee, Y.H.; Zhao, J.; Li, H.; Reyes, K. Machine-learning assisted exploration:
Toward the next-generation catalyst for hydrogen evolution reaction. J. Electrochem. Soc. 2021, 168, 126523. [CrossRef]
114. Hueffel, J.A.; Sperger, T.; Funes-Ardoiz, I.; Ward, J.S.; Rissanen, K.; Schoenebeck, F. Accelerated dinuclear palladium catalyst
identification through unsupervised machine learning. Science 2021, 374, 1134–1140. [CrossRef] [PubMed]
115. Zhang, J.; Zhu, Z.; Xiang, X.-D.; Zhang, K.; Huang, S.; Zhong, C.; Qiu, H.-J.; Hu, K.; Lin, X. Machine learning prediction of
superconducting critical temperature through the structural descriptor. J. Phys. Chem. C 2022, 126, 8922–8927. [CrossRef]
116. Le, T.D.; Noumeir, R.; Quach, H.L.; Kim, J.H.; Kim, J.H.; Kim, H.M. Critical temperature prediction for a superconductor: A
variational bayesian neural network approach. IEEE Trans. Appl. Supercond. 2020, 30, 8600105. [CrossRef]
117. Zhang, J.; Zhang, K.; Xu, S.; Li, Y.; Zhong, C.; Zhao, M.; Qiu, H.-J.; Qin, M.; Xiang, X.-D.; Hu, K. An integrated machine learning
model for accurate and robust prediction of superconducting critical temperature. J. Energy Chem. 2023, 78, 232–239. [CrossRef]
118. Roter, B.; Dordevic, S. Predicting new superconductors and their critical temperatures using machine learning. Phys. C Supercond.
Its Appl. 2020, 575, 1353689. [CrossRef]
119. Pereti, C.; Bernot, K.; Guizouarn, T.; Laufek, F.; Vymazalová, A.; Bindi, L.; Sessoli, R.; Fanelli, D. From individual elements to
macroscopic materials: In search of new superconductors via machine learning. NPJ Comput. Mater. 2023, 9, 71. [CrossRef]
120. Xie, Y.; Zhang, C.; Hu, X.; Zhang, C.; Kelley, S.P.; Atwood, J.L.; Lin, J. Machine learning assisted synthesis of metal–organic
nanocapsules. J. Am. Chem. Soc. 2019, 142, 1475–1481. [CrossRef]
Materials 2023, 16, 5977 30 of 30
121. Pellegrino, F.; Isopescu, R.; Pellutiè, L.; Sordello, F.; Rossi, A.M.; Ortel, E.; Martra, G.; Hodoroaba, V.-D.; Maurino, V. Machine
learning approach for elucidating and predicting the role of synthesis parameters on the shape and size of TiO2 nanoparticles.
Sci. Rep. 2020, 10, 18910. [CrossRef]
122. Tao, H.; Wu, T.; Aldeghi, M.; Wu, T.C.; Aspuru-Guzik, A.; Kumacheva, E. Nanoparticle synthesis assisted by machine learning.
Nat. Rev. Mater. 2021, 6, 701–716. [CrossRef]
123. Braham, E.J.; Cho, J.; Forlano, K.M.; Watson, D.F.; Arròyave, R.; Banerjee, S. Machine learning-directed navigation of synthetic
design space: A statistical learning approach to controlling the synthesis of perovskite halide nanoplatelets in the quantum-
confined regime. Chem. Mater. 2019, 31, 3281–3292. [CrossRef]
124. Epps, R.W.; Bowen, M.S.; Volk, A.A.; Abdel-Latif, K.; Han, S.; Reyes, K.G.; Amassian, A.; Abolhasani, M. Artificial chemist: An
autonomous quantum dot synthesis bot. Adv. Mater. 2020, 32, 2001626. [CrossRef] [PubMed]
125. Wang, J.; Wang, Y.; Chen, Y. Inverse design of materials by machine learning. Materials 2022, 15, 1811. [CrossRef] [PubMed]
126. Li, S.; Barnard, A.S. Inverse Design of Nanoparticles Using Multi-Target Machine Learning. Adv. Theory Simul. 2022, 5, 2100414.
[CrossRef]
127. Wang, R.; Liu, C.; Wei, Y.; Wu, P.; Su, Y.; Zhang, Z. Inverse design of metal nanoparticles based on deep learning. Results Opt.
2021, 5, 100134. [CrossRef]
128. He, J.; He, C.; Zheng, C.; Wang, Q.; Ye, J. Plasmonic nanoparticle simulations and inverse design using machine learning.
Nanoscale 2019, 11, 17444–17459. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual
author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to
people or property resulting from any ideas, methods, instructions or products referred to in the content.
ML can reduce computational costs, shorten development cycles, and improve computational accuracy, making it a promising approach for novel materials screening and predicting material properties. It provides guidance for stable and efficient material synthesis while offering novel ideas and research methodologies for materials science .
The ip-DNN model brought innovation to solubility prediction by using a multi-step approach, predicting intermediate molecular components and descriptors first, to enhance accuracy. It analyzed input data like NMR information, refractive index, and density, and used intermediate regression models to refine predictions, demonstrating a unique methodology in solubility prediction .
The major challenges include the requirement for large amounts of data, which are difficult to obtain due to high acquisition costs and lack of standardized data processing. Moreover, the datasets often suffer from excessive concentration or dispersion. The selection of appropriate features is critical but challenging, and manual feature engineering can overlook significant details, whereas automated feature engineering offers a possible solution. Additionally, ML models are often 'black boxes,' necessitating experimental verification and further study of the underlying physicochemical laws .
Liang et al. developed a generalized ML method using artificial neural networks (ANNs) to predict polymer compatibility, which is the total miscibility of polymers at the molecular scale. They built a database by aggregating data from scattered literature using natural language processing techniques. The method allowed predictions based on the molecular structure of the polymers and their compositions, achieving a prediction accuracy of at least 75% on a dataset with 1400 entries .
The use of "black-box" ML models implies a lack of interpretability, which necessitates experimental verification of results and a deep understanding of underlying physicochemical laws. While they offer potential in exploratory tasks and can guide synthesis and property prediction, traditional methods remain critical due to ML models' current limitations in explanation and understanding .
Machine learning accelerates the discovery of single-atom catalysts for nitrogen fixation by quickly analyzing vast datasets to identify promising catalyst candidates. This method leverages ML's ability to process and evaluate extensive materials data efficiently, bypassing the slow, traditional trial-and-error experimental approaches .
Automated feature engineering is beneficial because it can automatically construct new candidate features and select those most appropriate for model training, effectively addressing challenges where manual feature selection might be biased by the researcher's intuition or experience, potentially overlooking important features .
The atomic table CNN model predicts material properties such as the band gap and ground energy with greater accuracy compared to traditional DFT calculations. This model not only provides accurate results but does so with lower computational costs, highlighting ML's potential to outperform traditional methods in specific applications within material science .
The potential future of ML in materials science is tremendous, with prospects of eventually replacing traditional research methods by bridging exploratory tasks and offering innovative research strategies. Further theoretical and methodological improvements might enable ML to take on more comprehensive roles in materials prediction, discovery, and synthesis, beyond its current exploratory capacity .
Data preprocessing and feature engineering enhance the dataset's structure, enabling computers to better understand physicochemical relationships of materials, thereby improving the detection and prediction of material properties. High-quality data is crucial for the effective functioning of ML models, as the final results highly depend on the data's amount and reliability .