0% found this document useful (0 votes)
26 views30 pages

Machine Learning in Material Innovation

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
26 views30 pages

Machine Learning in Material Innovation

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

materials

Review
Application of Machine Learning in Material Synthesis and
Property Prediction
Guannan Huang, Yani Guo, Ye Chen and Zhengwei Nie *

School of Mechanical and Power Engineering, Nanjing Tech University, Nanjing 211816, China;
i35653184@[Link] (G.H.); guo17793470078@[Link] (Y.G.); chenye@[Link] (Y.C.)
* Correspondence: niez@[Link]

Abstract: Material innovation plays a very important role in technological progress and industrial
development. Traditional experimental exploration and numerical simulation often require con-
siderable time and resources. A new approach is urgently needed to accelerate the discovery and
exploration of new materials. Machine learning can greatly reduce computational costs, shorten the
development cycle, and improve computational accuracy. It has become one of the most promising
research approaches in the process of novel material screening and material property prediction.
In recent years, machine learning has been widely used in many fields of research, such as super-
conductivity, thermoelectrics, photovoltaics, catalysis, and high-entropy alloys. In this review, the
basic principles of machine learning are briefly outlined. Several commonly used algorithms in
machine learning models and their primary applications are then introduced. The research progress
of machine learning in predicting material properties and guiding material synthesis is discussed.
Finally, a future outlook on machine learning in the materials science field is presented.

Keywords: machine learning; material screening; property prediction; material synthesis;


artificial intelligence

1. Introduction
Citation: Huang, G.; Guo, Y.; Chen,
New materials have become the cornerstone of scientific and technological develop-
Y.; Nie, Z. Application of Machine
ment. Discovering materials with targeted properties, especially nanomaterials, has always
Learning in Material Synthesis and
Property Prediction. Materials 2023,
been a hotspot in science [1,2]. At present, the research and development of new materials
16, 5977. [Link]
mainly relies on researchers’ intuitive judgment of materials and empirical trial-and-error
10.3390/ma16175977
methods, which are not only inefficient but also often require a certain level of experience
and luck to obtain the target materials. At the same time, methods based on density func-
Academic Editor: Dorota
tional theory (DFT) are widely used in the research and development of novel materials.
Wilk-Kołodziejczyk
Since their initial development, DFT methods have evolved from limited calculations that
Received: 31 July 2023 provide approximate results to increasingly accurate and predictable methods. These meth-
Revised: 22 August 2023 ods have made important contributions in a variety of fields, such as materials discovery
Accepted: 28 August 2023 and design, drug design, solar cells, and hydrolytic materials [3]. The accuracy of these
Published: 31 August 2023 methods, however, is limited in fast calculations. To obtain high-accuracy results, the
computational volume often has to be much higher, which is difficult to exploit efficiently
in the research and development of new materials. In this context, artificial intelligence (AI)
is becoming highly popular with researchers as a means of accelerating the development of
Copyright: © 2023 by the authors. innovative materials. A subfield of AI that has grown rapidly in recent years is machine
Licensee MDPI, Basel, Switzerland.
learning (ML). ML applications are built on statistical algorithms. ML performs similarly
This article is an open access article
to researchers’ performance [4]. Because of its powerful data processing capability and
distributed under the terms and
relatively low research threshold, ML can effectively reduce human and material costs in
conditions of the Creative Commons
the process of novel material development and shorten the research and development cycle.
Attribution (CC BY) license (https://
By replacing or collaborating with traditional experiments and computational simulations,
[Link]/licenses/by/
4.0/).
ML could be employed to analyze material structures and predict material properties,

Materials 2023, 16, 5977. [Link] [Link]


Materials 2023, 16, x FOR PEER REVIEW 2 of 30
Materials 2023, 16, 5977 2 of 30

predict material properties, enabling the development of novel functional materials more
enabling the development of novel functional materials more efficiently and accurately.
efficiently and accurately. As a result, ML has become one of the most crucial methods for
As a result, ML has become one of the most crucial methods for replacing traditional
replacing traditional research and development. In the recent past, researchers in different
research and development. In the recent past, researchers in different fields, including
fields, including computer scientists and experts in AI algorithms, have used this ap-
computer scientists and experts in AI algorithms, have used this approach extensively,
proach contributing
greatly extensively, to
greatly contributingoftoML
the development thetechniques
development of ML
[5]. ML techniques
is now [5]. ML in
widely utilized is
now widely utilized in fields such as natural language understanding,
fields such as natural language understanding, non-monotonic reasoning, machine vision, non-monotonic
reasoning,
and patternmachine vision,
recognition [6]. and pattern recognition [6].
The basic
The basic principle
principle of ML is
of ML is to
to learn
learn (or
(or guess)
guess) general
general patterns
patterns from
from aa limited
limited
amount of training data and use these patterns to make predictions on unknown data.
amount of training data and use these patterns to make predictions on unknown data.
Figure 11 shows
Figure showsan anML
MLworkflow
workflow example.
example. MLML hashasbeenbeen
usedused
to detect the solubility
to detect of C60
the solubility
in materials
of science as
C60 in materials early as
science as the lastascentury
early the last[7]. It is now
century [7].used
It isto
nowdiscover
used novel mate-
to discover
rials, predict material and molecular properties, study quantum
novel materials, predict material and molecular properties, study quantum chemistry,chemistry, and design
drugs.
and The drugs.
design purposeTheof purpose
this review is toreview
of this offer an overview
is to offer an of the employment
overview of ML in
of the employment
predicting material properties and performance, guiding material
of ML in predicting material properties and performance, guiding material synthesis, synthesis, and project-
ing models and conclusions. This review not only provides guidance
and projecting models and conclusions. This review not only provides guidance for for researchers to
synthesize stable and efficient materials, but also inspires their interest in
researchers to synthesize stable and efficient materials, but also inspires their interest in the use of ML
in materials
the use of MLresearch.
in materials research.

workflow.
Figure 1. An example of an ML workflow.

2. Data Pre-Processing
2. Data Pre-Processing
If
If ML models
ML models are
are the
the engines
engines that
that handle
handle various
various tasks,
tasks, data
data are
are the
the fuel
fuel that
that drives
drives
the models. A sufficient amount of data is a prerequisite to making the model
the models. A sufficient amount of data is a prerequisite to making the model work. High- work.
High-quality
quality data enable the model to run effectively. Due to this, large amounts of datadata
data enable the model to run effectively. Due to this, large amounts of are
are critical
critical to ML
to ML [8]. [8]. In general,
In general, the final
the final ML ML results
results are directly
are directly affected
affected by amount
by the the amount
and
and reliability of the data. This is where data pre-processing and feature engineering are
reliability of the data. This is where data pre-processing and feature engineering are ben-
beneficial. Data pre-processing and feature engineering could promote the reconstruction of
eficial. Data pre-processing and feature engineering could promote the reconstruction of
datasets so that computers could more easily understand the physicochemical relationships
datasets so that computers could more easily understand the physicochemical relation-
of materials, detect material properties, and build prediction models [9].
ships of materials, detect material properties, and build prediction models [9].
2.1. Data Collection and Cleaning
2.1. Data Collection and Cleaning
2.1.1. Data Collection
[Link]
Data
ML,Collection
the size and quality of the training dataset employed for learning could
In ML, the
significantly sizethe
affect and quality of
accuracy of athe training model.
predictive dataset Therefore,
employed training
for learning couldneed
datasets sig-
nificantly
to affectorthe
be collected accuracy
created of a predictive
carefully. In general, model. Therefore,
training data cantraining datasets
be gathered needways.
in three to be
collected ordata
Obtaining created
fromcarefully. In general,
the published training
literature is thedata
firstcan be gathered
method. in three
The data ways.
obtained Ob-
in this
taining
way databefrom
could morethe published
relevant and literature
provide aisdirection
the first for
method. The data
synthesis obtained in [10].
and application this
way could
Second, be more relevant
high-throughput and provide or
computations a direction
experiments for synthesis andtoapplication
can be used obtain data.[10].
It
Second,behigh-throughput
should noted that, in some computations
cases, theseordata experiments can be used
may be incomplete, to obtain or
inconsistent, data.
even It
should be[11].
spurious notedThe that,third
in some cases,
method is these data may
to obtain data be incomplete,
from inconsistent,
open databases or even
available on
repository websites.
spurious [11]. The third The Materials
method is toGenome
obtain data Initiative,
from open initiated by the
databases Unitedon
available States
repos-in
2011,
itory emphasizes
websites. The theMaterials
importance of massive
Genome data in
Initiative, the development
initiated of materials
by the United science,
States in 2011,
which encourages
emphasizes the development
the importance of massive of high-quality material databases
data in the development [12]. With
of materials the
science,
continuous development of theoretical and experimental research,
which encourages the development of high-quality material databases [12]. With the con- data generated from
experiments and computational
tinuous development simulations,
of theoretical including failure
and experimental data,data
research, have generated
been integrated
from
Materials 2023, 16, 5977 3 of 30

into databases [13]. These databases are based on the concept of material data sharing,
which greatly simplifies the process of obtaining material information. Table 1 introduces
some commonly used methods for collecting data from publicly available databases.
For instance, Zhou et al. [14] developed an ML-based approach to predict cathode
materials for Zn-ion batteries with high capacity and high voltage. They screened over
130,000 inorganic materials from the materials project database and applied a crystal
graph convolutional-neural-network-based ML approach with data from the Automatic
Flow (AFLOW) database. This resulted in the prediction of approximately 80 cathode
materials, with 10 of them being experimentally discovered previously and agreeing
well with the observed measurements. Additionally, approximately 70 new promising
candidates were predicted for further experimental validation.

Table 1. An overview of some databases in material science.

Database Website Brief Introduction


[Link] A globally available database of 3,530,330 material compounds with
AFLOW
(accessed on 17 July 2023) over 734,308,640 calculated properties and growing.
Crystallography Open [Link] Open-access collection of crystal structures of organic, inorganic,
Database (COD) (accessed on 17 July 2023) metal–organic compounds and minerals, excluding biopolymers.
The world’s largest database of small-molecule organic and
Cambridge Structural [Link]
metal–organic crystal structure data, now at over
Database (CSD) (accessed on 17 July 2023)
1.2 million structures.
A comprehensive collection of crystal structure information for
non-organic compounds, including inorganics, ceramics, minerals,
Inorganic Crystal Structure [Link]
and metals, covers the literature from 1915 to the present and
Database (ICSD) (accessed on 17 July 2023)
contains over 60,000 entries on the crystal structure
of in-organic materials.
[Link] A database containing 154,718 materials, 4351 intercalation
Materials Project
(accessed on 17 July 2023) electrodes, and 172,874 molecules.
Open Quantum Materials [Link] The OQMD is a database of DFT calculated thermodynamic and
Database (OQMD) (accessed on 17 July 2023) structural properties of 1,022,603 materials.

2.1.2. Data Cleaning


When collecting raw data, unprocessed datasets are difficult to analyze and some-
times become useless, as they tend to be inconsistent, missing, and noisy. Before using
those datasets, quality must be maintained. Data cleaning is an operation performed
on the existing data to remove anomalies and obtain the data collection, which is an
accurate and unique representation of the mini world. It involves eliminating errors,
resolving inconsistencies, and transforming the data into a uniform format [15]. Data
cleaning is an enormous task achieved by smoothing noise, completing missing values,
correcting inconsistencies, and identifying outliers in data. The common methods for
filling in missing values are as follows: fill in missing values manually; fill in missing
values with a global constant; fill in missing values with the average value of attributes;
fill in corresponding missing values with the average value of attributes of the same type
as the given tuple; and fill in missing values with the most likely value. The commonly
used methods for smoothing noise are binning, regression, and clustering [10]. Binning
is employed to handle noisy data. In this approach, the data are sorted, and then values
are partitioned by equal-frequency bins where values are put into an equal number of
bins. Regression involves predicting unknown data from known data and fitting it using
a function. The two types of regression techniques are linear and multiple linear. Linear
regression uses a known value to predict an unknown value, fitting the relationship
between the two values with a straight line. To reduce outliers, clustering can be imple-
mented. Clustering refers to grouping data points with similar properties into clusters.
By categorizing outliers as points outside these clusters, they could be easily identi-
fied and minimized in the dataset [16–18]. Data cleaning can effectively improve the
model’s prediction accuracy. Liu et al. [19] discussed the prediction of protein–protein
known value to predict an unknown value, fitting the relationship between the two values
with a straight line. To reduce outliers, clustering can be implemented. Clustering refers
to grouping data points with similar properties into clusters. By categorizing outliers as
points outside these clusters, they could be easily identified and minimized in the dataset
Materials 2023, 16, 5977 [16–18]. Data cleaning can effectively improve the model’s prediction accuracy. Liu 4etofal. 30
[19] discussed the prediction of protein–protein interaction sites using ML-based compu-
tational approaches. The authors proposed a method that improves prediction perfor-
mance by addressing
interaction sites usingthe class imbalance
ML-based issue in protein–protein
computational approaches. The interaction site predic-a
authors proposed
tion.
methodThey operated
that improves a data-cleaning procedure by
prediction performance to remove
addressing marginal targets
the class from majority
imbalance issue in
samples and a post-filtering
protein–protein procedure
interaction site to reduce
prediction. Theyfalse-positive predictions. The
operated a data-cleaning proposed
procedure to
remove marginal
method was tested targets from majority
on benchmark samples
datasets and and a post-filtering
showed competitiveprocedure
performance to reduce
com-
false-positive
pared predictions.
to existing predictors. The proposed method was tested on benchmark datasets and
showed competitive performance compared to existing predictors.
2.2. Feature Engineering
2.2. Feature Engineering
A key part of the data preparation phase in ML is feature engineering. It extracts
A key
features part
(also of theasdata
known preparation
descriptors) fromphase
the rawin data
ML is andfeature engineering.
transforms It extracts
the features into a
features (also known as descriptors) from the raw data and transforms
format suitable for ML models. The selection of features is critical for building ML models the features into a
format suitable for ML models. The selection of features is critical for
and could even determine the upper limit of overall model performance [20]. In feature building ML models
and coulddifferent
selection, even determine
parametersthe upper
could limit of overall
be operated as model
features performance
for chemical [20].
andInmaterial
feature
selection, different parameters could be operated as features for chemical
structures (and their properties), e.g., electronic properties (band gap, dielectric constant, and material
structures
work (andelectron
function, their properties), e.g.,electron
density, and electronic properties
affinity) (bandfeatures
and crystal gap, dielectric constant,
(translation vec-
work function, electron density, and electron affinity) and crystal features
tors, fractional coordinates of atoms, radial distribution functions, and Voronoi tessella- (translation vec-
tors, fractional coordinates of atoms, radial distribution functions, and Voronoi tessellations
tions of atomic positions). It is worth noting that rational feature selection is often expen-
of atomic positions). It is worth noting that rational feature selection is often expensive
sive and difficult [11]. In past studies, feature selection has typically had to be performed
and difficult [11]. In past studies, feature selection has typically had to be performed
manually. However, the limitations of manual feature engineering prevented the selection
manually. However, the limitations of manual feature engineering prevented the selection
of the most representative features in most cases. Over the last few years, the employment
of the most representative features in most cases. Over the last few years, the employment
of automated feature engineering has become increasingly widespread. It automatically
of automated feature engineering has become increasingly widespread. It automatically
constructs brand new candidate features from data and selects the most suitable features
constructs brand new candidate features from data and selects the most suitable features
for model training, which could solve the dilemma faced by manual feature engineering.
for model training, which could solve the dilemma faced by manual feature engineering.
Wang et al. [21] utilized automated feature engineering for the development of nano-
Wang et al. [21] utilized automated feature engineering for the development of nano-
materials. Automated feature engineering uses deep learning algorithms to automatically
materials. Automated feature engineering uses deep learning algorithms to automatically
develop
develop aa set of features
set of features that
that are
are relevant
relevant to to the
the desired
desired output.
output. As As aa result,
result, non-experts
non-experts
could
could select features much more easily, which would greatly reduce the use of of
select features much more easily, which would greatly reduce the use expertise
expertise in
in training models. The variation in feature engineering in the design
training models. The variation in feature engineering in the design of nanomaterials can be of nanomaterials
can be observed
observed in Figurein Figure
2. 2.

Figure 2. Evolution of the ML workflow in nanomaterial discovery and design. (a) First-generation
approach. In this paradigm, there are two main steps: feature engineering from raw database to de-
scriptors and model building from descriptors to target model. (b) Second-generation approach. The
key characteristic that distinguishes this approach from the first-generation approach is eliminating
human-expert feature engineering, which can directly learn from raw nanomaterials. Reproduced
with permission from [21].
Materials 2023, 16, 5977 5 of 30

3. Classification of ML and Algorithms


Once sufficient training data are selected, models can be built for the development of
novel materials. Choosing an appropriate algorithm for a training model is essential for
making accurate predictions. Based on the type of processed data, ML can be classified as
supervised learning, unsupervised learning, semi-supervised learning, and reinforcement
learning. For supervised learning, the input training data are labeled. After optimizing
the model with ML, a predictable output value for a new input value could be acquired.
In contrast, the input training data are unlabeled in unsupervised learning. Using an
algorithm, the unlabeled training set is trained to find potential features. As for semi-
supervised learning, the input training data are partially labeled. Reinforcement learning
occurs when the training object interacts with the environment, obtaining feedback from
the environment and adjusting its strategy to accomplish a specific goal or to maximize
the benefit of a behavior [22]. Next, a brief description of several commonly utilized ML
algorithms is given.

3.1. Shallow Learning


Shallow learning usually has no hidden layer or only one hidden layer [23]. The
approaches include decision tree (DT), K-nearest neighbor (KNN), support vector machine
(SVM) [24], random forest (RF), and artificial neural network (ANN). Shallow learning has
produced satisfactory results in various areas of materials science. In this section, some
algorithms for shallow learning are presented, some applications in materials science are
summarized, and the ML model used by the researchers is demonstrated.

3.1.1. KNN
The KNN algorithm was first proposed by Cover and Hart [25]. The KNN classification
is one of the most basic and simplest classification methods. It should be considered for
classification studies when little or no data distribution experience is available [26]. The
principle of the KNN algorithm is that if most of the most similar K samples in the feature
space (i.e., the nearest samples in the feature space) belong to a certain category, the sample
also belongs to this category. Figure 3 shows a schematic of a typical KNN algorithm. For
an unknown target, when K takes 3, the target is classified into class 1; when K takes 7,
the target is classified into class 2. According to this method, the sample’s category is
determined by its proximity to one or more nearby samples. The KNN algorithm itself
is simple and effective, easy to understand, and straightforward to implement. Since it
does not require prediction parameters or training, the KNN algorithm is suitable for time
classifications, especially for multimodals (i.e., objects with multiple categories). Recently,
KNN algorithms have been widely utilized in text classification, pattern recognition, image
processing, and materials science. Sharma et al. [27] employed the KNN algorithm to
predict the dynamic fracture toughness of glass-filled polymer composites. The dynamic
modulus of elasticity, aspect ratio, and volume fraction of glass particles were used as
independent model parameters. The proposed KNN model predicted the fracture behavior
of the composites with an accuracy of 96%. It is also possible to extend their model to
predict other material properties.
The drawback of the KNN algorithm is that as the amount of data increase, the
computational complexity of the KNN increases accordingly. This is because the KNN
algorithm needs to calculate both training data and test data for each classification or
regression. If there are a large amount of data, the computing power required would be
greatly increased. In addition, the randomness of training data also affects the performance
of the KNN algorithm [28].
Materials 2023, 16, x FOR PEER REVIEW 6 o
Materials 2023, 16, 5977 6 of 30

Figure 3. Schematic of a typical KNN algorithm.

The drawback of the KNN algorithm is that as the amount of data increase, the c
putational complexity of the KNN increases accordingly. This is because the KNN a
rithm needs to calculate both training data and test data for each classification or reg
sion. If there are a large amount of data, the computing power required would be gr
increased. In addition, the randomness of training data also affects the performance o
KNN algorithm [28].

3.1.2. DT
A DT is a typical classification method. The earliest DT algorithm was the con
Figure 3. Schematic of a typical KNN algorithm.
Figure 3. Schematic
learning of a typicalby
system proposed KNN algorithm.
Hunt [29]. The most influential DT algorithms are ID3
3.1.2. DT
and C4.5 [31], which were proposed by Quinlan in 1986 and 1993, respectively. DTs
The
sify A DTdrawback
is a typical
training data by of the KNN algorithm
classification
different method.
features,The isearliest
that as
aiming the
toDT amountcategorize
algorithm
correctly of data
was increase,
the concept the A
instances. c
learning
putational system proposed
complexity by Hunt [29].
of thedecision The
KNN increases most influential
accordingly. DT algorithms
This are
is because ID3 [30]
model consists of internal nodes and leaf nodes. Each internal the nodeKNNsplita
and
rithmC4.5
needs[31], which were
to calculate proposed by Quinlan in 1986 and 1993, respectively. DTs
instance space into two both or moretraining data and
subspaces test data
according to for each classification
a certain
classify training data by different features, aiming to correctly categorize instances. A
discrete function or rego
sion.
input If there
attribute are a large amount of data, the computing power required would be gre
DT model consistsvalues, anddecision
of internal each leaf node
nodes and isleaf
assigned
nodes. toEachone class representing
internal node splits the
increased.
appropriate
the In
instance space addition,
target
intovalue the randomness
two or[32].
moreChen of training
et [Link]
subspaces data
[11] presented also affects
thediscrete
to a certain structurethe performance
of a typical
function of of
D
KNN
shown
the algorithm
input Figure[28].
inattribute 4. A typical
values, and eachdecision treeis algorithm
leaf node assigned toconsists
one classof three mainthe
representing steps: fea
most appropriate target value [32]. Chen et al. [11] presented
selection, decision tree generation, and pruning. The purpose of pruning is to minithe structure of a typical
DT, asDT
3.1.2. shown in Figure 4. A typical decision tree algorithm consists of three main steps:
the structural risk of the model by optimizing the loss function and weighing the mo
feature selection, decision tree generation, and pruning. The purpose of pruning is to
A
minimize DT
complexity is
the anda typical
accuracy.
structural classification
risk ofLiu
the et method.
al. [33]
model developed
by optimizingThe earliest
a DT
the DT algorithm
lossmodel
function forand was the
predicting
weighing the con
resi
learning
tensile system
strength proposed
and modulus by Hunt
of [29]. The most influential
pultruded-fiber-reinforced
the model’s complexity and accuracy. Liu et al. [33] developed a DT model for predicting DT
polymeralgorithms
(FRP) are ID3
compo
and
Using
the C4.5an[31],
residual which
existing
tensile wereand
database,
strength proposed
746
modulusdataby Quinlan
points
of were in collected
1986 and for
pultruded-fiber-reinforced 1993, respectively.
training.
polymer DTs c
The accurac
(FRP)
composites.
sify trainingwas
the model Using
data an existing
by different
verified database, 746 data
features, The
experimentally. points
aiming were collected
to correctly
significance for training.
of allcategorize The
attributes instances.
of the input A
accuracy
model
was also ofquantitatively
the model
consists was analyzed
of internal verified
decisionexperimentally.
bynodes andThe
the model. leaf
Thesignificance
nodes.
proposed ofDT
Each allinternal
attributes
model node of splits
provides a
the input data was also quantitatively analyzed by the model. The proposed DT model
instance
method space into twothe
for predicting or long-term
more subspaces according
degradation of FRPto acomposites
certain discrete subjectedfunction
to envof
provides a new method for predicting the long-term degradation of FRP composites
input
mental attribute
subjected influences. values, and
to environmental each leaf node is assigned to one class representing the m
influences.
appropriate target value [32]. Chen et al. [11] presented the structure of a typical DT
shown in Figure 4. A typical decision tree algorithm consists of three main steps: fea
selection, decision tree generation, and pruning. The purpose of pruning is to minim
the structural risk of the model by optimizing the loss function and weighing the mod
complexity and accuracy. Liu et al. [33] developed a DT model for predicting the resid
tensile strength and modulus of pultruded-fiber-reinforced polymer (FRP) compos
Using an existing database, 746 data points were collected for training. The accurac
the model was verified experimentally. The significance of all attributes of the input d
was also quantitatively analyzed by the model. The proposed DT model provides a n
method for predicting the long-term degradation of FRP composites subjected to envi
mental influences.

Figure 4. Diagram of a DT. The circles and squares indicate internal nodes and leaf nodes, respectively.
Different colors represent different classes. Reproduced with permission from [11].
Figure 4. Diagram of a DT. The circles and squares indicate internal nodes and leaf nodes, res
tively. Different colors represent different classes. Reproduced with permission from [11].

The RF algorithm consists of multiple DTs. In RFs, each tree casts a unit vote for
Materials 2023, 16, 5977 7 of 30
most popular class, and then combining these votes obtains the final sort result. RFs
sess high classification accuracy [34]. It would, however, take a great deal of space
timeThe
to RF
train an RF consists
algorithm with many DTs. Compared
of multiple DTs. In RFs,with
each DTs, the acalculation
tree casts unit vote forcosts of
would also increase significantly. In this regard, RFs and DTs should be selected
the most popular class, and then combining these votes obtains the final sort result. RFs base
the actual
possess situation.
high classification accuracy [34]. It would, however, take a great deal of space
and time to train an RF with many DTs. Compared with DTs, the calculation costs of RFs
would also increase significantly. In this regard, RFs and DTs should be selected based
3.1.3. ANN
on the actual situation.
The concept of an ANN was introduced by McCulloch and Pitts [35]. An ANN
3.1.3. ANNnetwork structure that is formed by a large number of nodes (neurons) conne
complex
to each
Theother.
concept Itof
is an
a kind
ANNof abstraction,
was introduced simplification,
by McCulloch and and simulation
Pitts [35]. An ANN of the
is organiza
a
complex network structure that is formed by a large number of nodes
and operation mechanism of the human brain. Each node in an ANN represents a spe (neurons) connected
to each other. It is a kind of abstraction, simplification, and simulation of the organization
output function, i.e., the activation function. Each connection between any two nodes
and operation mechanism of the human brain. Each node in an ANN represents a specific
resents a weighted
output function, value
i.e., the for thefunction.
activation signal passing through that
Each connection connection,
between which is equ
any two nodes
lent to theamemory
represents weightedofvaluethe [Link] theThe network’s
signal connection
passing through that mode, the value
connection, whichofisthe weig
and the excitation
equivalent function
to the memory all have
of the [Link] Theeffect on its connection
network’s output [36]. mode,As athemajor soft-compu
value of
the weights, and
technology, ANNsthe excitation
have been function all have
extensively an effectand
studied on its outputin
applied [36]. As adecades
recent major [37].
soft-computing technology,
The structure ANNsANN
of a typical have been extensively
is shown studied
in Figure andnodes
5. Its appliedareingenerally
recent divi
decades [37].
into three categories: input, hidden, and output. The input nodes represent the in
The structure of a typical ANN is shown in Figure 5. Its nodes are generally divided
mation
into threereceived
categories:from
input,the inputand
hidden, data. TheThe
output. output
inputnodes are utilized
nodes represent to store the resul
the information
the datafrom
received processing.
the inputThe [Link] between
The output the are
nodes input and output
utilized to store nodes are of
the results so-called
the hid
nodes. DifferentThe
data processing. types
nodesof between
nodes inthe aninput
ANNand areoutput
distributed
nodes arein multiple layers. The no
so-called hidden
nodes.
on Different
different typescould
layers of nodes be in an ANN are
connected bydistributed
lines, whichin multiple
[Link] nodes in ne
synapses
on different layers could be connected by lines, which correspond
structures, representing a nonlinear mapping. The learning process of an ANN is to to synapses in neural
structures, representing a nonlinear mapping. The learning process of an ANN is to
tinuously optimize the whole network model by correcting the weights of nodes in e
continuously optimize the whole network model by correcting the weights of nodes in each
layer withtraining
layer with trainingdatadata
[38].[38].

Figure [Link]
Figure Diagramof aoftypical ANN.
a typical ANN.
A variety of ANN models and their variants have been developed. The variants
A variety of ANN models and their variants have been developed. The variant
include back-propagation networks, perceptrons, self-organizing mappings, Hopfield
clude back-propagation
networks, networks,ANNs
and Boltzmann machines. perceptrons,
have beenself-organizing
applied to drivemappings, Hopfield
the synthesis
works,
of a wide range of functional materials, such as shape memory alloys [39], hyperelastic of a w
and Boltzmann machines. ANNs have been applied to drive the synthesis
materials
range of [40], and high-entropy
functional materials,alloys
such(HEAs) [41]. memory
as shape Table 2 illustrates the application
alloys [39], of mate
hyperelastic
the afore-mentioned algorithms.
[40], and high-entropy alloys (HEAs) [41]. Table 2 illustrates the application of the af
mentioned algorithms.

Table 2. Some applications of shallow learning in materials science.


Materials 2023, 16, 5977 8 of 30

Table 2. Some applications of shallow learning in materials science.

Researchers
Materials 2023, 16, x FOR PEER REVIEW Algorithms Purposes 8 of 30
Predict the fracture toughness of silica-filled
Sharma et al. [42] KNN
epoxy composites.
Predict surface roughness in the micro-plasma
Researchers Algorithms Purposes
Kumar et al. [43] KNN transfer arc metal additive manufacturing
Sharma et al. [42] KNN Predict the fracture toughness (µ-PTAMAM)
of silica-filledprocess.
epoxy composites.
Predict surface roughness in the micro-plasma transfer arc metal
KumarJalali
et [Link][43]
al. [44] KNN KNN (Figure 6a) Predict phases in HEAs.
additive manufacturing (µ-PTAMAM) process.
Achieve rapid detection of transformer
Jalali Wang et al. [45]
et al. [44] KNN (Figure 6a) SVM Predict phases in HEAs.
winding materials.
Wang et al. [45] SVM Achieve rapid detection of transformer winding materials.
Predict the fracture life of martensitic steels under
Martinez et al. [46] SVM and ANN the fracture life of martensitic steels under high-tempera-
Predict
Martinez et al. [46] SVM and ANN high-temperature creep conditions.
ture creep conditions.
Adaptive boosting, RF, and DT Predict the compressive strength of concrete
Ahmad et al. [47] Adaptive boosting, RF,
Ahmad et al. [47] (Figure 6b) the compressive strengthatofhigh
Predict temperatures.
concrete at high temperatures.
and DT (Figure 6b)
Gradient boosted regression tree
Sun et al. [48] Gradient boosted regres- Evaluate the strength of coal–grout materials.
(GBRT) and RF
Sun et al. [48] Evaluate the strength of coal–grout materials.
sion tree (GBRT) and RF Predict the higher heating value (HHV) of biomass
Samadia et al. [49] GBRT
Predict the higher heating value (HHV)
materials based onof proximate
biomass materials
analysis. based
Samadia et al. [49] GBRT
on proximate
Predict analysis.
the compressive strength of eco-friendly
Shahmansouri et al. [50] Predict
ANN (Figure 6c)the compressive strength
geopolymer of eco-friendly
concrete incorporating geopolymer con-
silica fume and
Shahmansouri et al. [50] ANN (Figure 6c) natural zeolite.
crete incorporating silica fume and natural zeolite.
Development of a predictive
Development model for the chloride
of a predictive diffusion
model for coef-
the chloride
Liu etLiu
al. et al. [51]
[51] ANN ANN
diffusion coefficient in concrete.
ficient in concrete.

Figure 6. (a) A portion of the HEA interaction network with Fruchterman Reingold layout, adapted
with permission
with permission from
from [44].
[44]. (b)
(b) Schematic
Schematic illustration
illustration of
of an
an RF
RF structure,
structure, adapted
adapted with
with permission
permission
from [47]. (c) A multi-layer neural network model layout, adapted with permission from
from [47]. (c) A multi-layer neural network model layout, adapted with permission from [50].[50].

3.2. Deep Learning


Hinton et al. [52] first proposed the concept of deep learning. The unsupervised
greedy training layer-by-layer algorithm based on deep degree nets was designed to solve
optimization problems related to deep structures. Similar to an ANN, deep learning is a
multilayer neural network [53].
Materials 2023, 16, 5977 9 of 30

3.2. Deep Learning


Hinton et al. [52] first proposed the concept of deep learning. The unsupervised
greedy training layer-by-layer algorithm based on deep degree nets was designed to solve
optimization problems related to deep structures. Similar to an ANN, deep learning is a
multilayer neural network [53].

3.2.1. Overview of Deep Learning


Deep learning can be considered a subset of ML. The idea of deep learning is derived
from multilayer ANNs. The learning process of deep learning exhibits depth to some extent
because of the multilayer structure of ANNs. In each hidden layer, neurons receive input
signals from other neurons, combine them with their internal state, and produce output
signals. The connections between neurons have weights assigned to them, forming the
overall layer of a neural network. The learning process involves adapting the network by
adjusting the weights of the connections to minimize output errors. Deep learning, with
its self-adapting architecture, reduces the need for feature engineering and could identify
and work around defects that may be difficult to detect in other techniques [5]. Instead, the
algorithm adjusts itself in continuous learning and independently selects suitable features.
This could be viewed as a major advancement in ML. While traditional ML models may
be more accurate with small data, deep learning models tend to be more reliable when
big data is available. Deep neural networks (DNNs) with multiple hidden layers have
higher learning capacity, allowing them to saturate accuracy gains compared to traditional
models. Although training neural networks is computationally expensive, once trained,
deep learning can make very fast predictions. This one-time training cost is outweighed
by the speed of subsequent predictions [54]. After years of development, a variety of
deep learning models have been produced, mainly including stacked autoencoders [55],
deep belief networks (DBNs) [56], deep Boltzmann machines (DBMs) [57], DNNs [58], and
convolutional neural networks (CNNs) [59]. Deep learning techniques are widely utilized
in speech recognition, visual object recognition, object detection, drug discovery, and
genomics [60]. They are also some of the fastest-growing and most adaptable techniques
ever developed in materials science.
Additionally, deep learning faces the dilemma of how to effectively process large
amounts of complex data. In practical applications, building suitable deep learning models
is increasingly challenging. Although deep learning is not yet fully mature and has many
problems to solve, it has shown a strong learning capability. Throughout the future, deep
learning is expected to remain a key research focus in AI.

3.2.2. Applications of Deep Learning


Deep learning has been widely applied in materials science due to its excellent perfor-
mance. Based on industrial data, Wu et al. [61] investigated the impact energy prediction
model of low-carbon steel. A three-layer neural network, extreme learning machine, and
DNN were compared with different activation functions, structure parameters, and training
functions. Bayesian optimization was employed to determine the optimal hyper-parameters
of the DNN. The model with the highest performance was applied to investigate the im-
portance of process parameter variables on the impact energy of low-carbon steel. The
results showed that the DNN obtained better prediction results than those of a shallow
neural network because the multiple hidden layers improved the learning ability of the
model. Sun et al. [62] applied deep learning to rapidly predict the photovoltaic properties
of organic photovoltaic materials, with a prediction accuracy up to 91%. Konno et al. [63]
reported a deep learning algorithm for discovering novel superconductors. The prediction
accuracy of their ML model for material superconductivity was as high as 62%. Employing
the ML model, the authors found two superconductors that were not in the database and
found Fe-based high-temperature superconductors (discovered in 2008) in the training
data before 2008. These results pave the way for the discovery of new high-temperature
superconductors. Li et al. [64] explored a correlated deep learning framework consisting of
Materials 2023, 16, 5977 10 of 30

three recurrent neural networks (RNNs) to efficiently generate new energetic molecules
with high detonation velocity in the low data regime. They utilized data augmentation by
fragment shuffling of 303 energetic compounds to pretrain the RNN and then fine-tuned
it using the 303 compounds to produce molecules similar to the energetic compounds.
They also employed a simplified molecular input line entry (SMILE) system coupled with
pretrained knowledge to build an RNN-based prediction model for screening molecules
with high detonation velocity. Their strategy performed comparably to transfer learning
based on an existing big database. Quantum mechanics calculations confirmed that 35 new
molecules have higher detonation velocity and lower synthetic accessibility than the classic
explosive hexogen, with three novel molecules comparable to caged China Lake Compound
No. 20 in detonation velocity. Zhang et al. [65] utilized generative adversarial networks
(GANs) to design metaporous materials for sound absorption (Figure 7a). The researchers
trained the GANs using numerically prepared data and successfully developed designs
with high-standard broadband absorption performance. The GANs accelerated the design
process by hundreds of times, allowing for instantaneous multiple solutions. The GANs
also demonstrated the ability to generate creative configurations and rich local features.
This work highlighted the potential of ML in guiding the design and optimization process
for materials and opened up new possibilities for interdisciplinary research in AI and
materials. Unni et al. [66] introduced a deep convolutional mixture density network (MDN)
approach for the inverse design of layered photonic structures. The MDN modeled the
design parameters as multimodal probability distributions, allowing for convergence in
cases of nonuniqueness without sacrificing degenerate solutions. The MDN was applied to
the inverse design of two types of multilayer photonic structures consisting of thin films
of oxides, which present a challenge for conventional machine learning algorithms due
to their large degree of nonuniqueness in their optical properties. The MDN can handle
the transmission spectra of high complexity and varying illumination conditions. The
shape of the probability distributions provides valuable information for postprocessing
and prediction uncertainty. The MDN approach offers an effective solution to the inverse
design of photonic structures with high degeneracy and spectral complexity.
The use of vision transformers, residual networks (ResNets), and region-based-CNNs
(R-CNNs) on materials datasets has shown exceptional performance. Huang et al. [67]
proposed a waste materials classification method based on a vision transformer model
(Figure 7b). The model overcame CNN limitations by using self-attention mechanisms
to allocate weights to different parts of waste images. The vision transformer achieved
an accuracy rate of 96.98% by pretraining on ImageNet and fine-tuning on the TrashNet
dataset. The trained model can be deployed on a cloud server and accessed through
a portable device for real-time waste classification, which is convenient and efficient for
resource conservation and recycling. Jiang et al. [68] explored the use of global optimization
networks (GLOnets) with the ResNet architecture for the multiobjective and categorical
global optimization of photonic devices. The authors demonstrated that these networks,
called Res-GLOnets, could be configured to design thin-film stacks consisting of multiple
material types. The Res-GLOnets can find the global optimum with faster speeds compared
to conventional algorithms. The authors also showed the utility of their method for complex
design tasks, such as designing incandescent light filters. Wang et al. [69] proposed an
image detection method based on an improved Faster R-CNN model for wear location
and wear mechanism identification (Figure 7c). They trained and tested the model using
a wear image dataset produced by a self-made tribometer equipped with an imaging
system. The results showed that the proposed method had a detection accuracy of
more than 99%. It outperformed edge detection technology and Yolov3 target detection
models in wear location and wear mechanism identification. This research contributes
to the development of an innovative approach for the online and intelligent wear status
detection of machinery components.
Materials 2023, 16, x FOR PEER REVIEW 11 of 30
Materials 2023, 16, 5977 11 of 30

[Link]
Figure Some deep
deep learning
learning algorithm
algorithm structures.
structures. (a) Schematic
(a) Schematic illustration
illustration of procedures
of the design the design proce-
dures of metaporous materials with GANs, adapted with permission from [65].
of metaporous materials with GANs, adapted with permission from [65]. (b) Structure of a vision (b) Structure
trans- of a
vision transformer, adapted with permission from [67]. (c) Illustration of the concept of
former, adapted with permission from [67]. (c) Illustration of the concept of using image identification using image
identification
based based on
on the improved theR-CNN
Faster improved Faster
model R-CNN
to identify model
wear, to identify
adapted wear, adapted
with permission from [69].with permis-
sion from [69].
3.3. Materials Informatics Based on ML
3.3. Materials
Materials informatics
InformaticsisBased on ML
a study field that focuses on investigating and applying infor-
maticsMaterials
techniquesinformatics
to materialsisscience
a studyand engineering.
field Propelled
that focuses partly by theand
on investigating Materials
applying in-
Genome Initiative and partly by algorithmic developments and successes of
formatics techniques to materials science and engineering. Propelled partly by the Mate- data-driven
efforts in other domains,
rials Genome Initiative informatics strategies
and partly are beginning
by algorithmic to take shapeand
developments within materialsof data-
successes
science. Informatics strategies give rise to surrogate ML methods that can realize accurate
driven efforts in other domains, informatics strategies are beginning to take shape within
prediction using just historical data instead of experiments or simulations/calculations.
materials science. Informatics strategies give rise to surrogate ML methods that can realize
This methodology is usually composed of three distinct steps: acquisition of reliable
accurate data,
historical prediction using
statistical just historical
quantification data instead of material
of information-rich experiments or simulations/calcu-
structures, and map-
lations. This methodology is usually composed of three
ping between “input” and “output”. The commonly used ML algorithms in distinct steps: acquisition
materials of reli-
able historical
informatics data,
include statistical
regression, quantification
DT, ANN, and deep of learning
information-rich
[70–73]. Tomaterial
meet thestructures,
require- and
mapping
ments between
of the studies“input” and “output”.
of computational The informatics,
materials commonly used Zhao MLet [Link]
[74] derivedin an
materials
artificial-intelligence-aided data-driven infrastructure called Jilin Artificial-intelligence
informatics include regression, DT, ANN, and deep learning [70–73]. To meet the require-
aided
mentsMaterials-design
of the studies of Integrated Packagematerials
computational (JAMIP). informatics,
The organizationZhaoofetJAMIP abides
al. [74] derived an
by the data lifecycle in computational materials informatics, from data generation
artificial-intelligence-aided data-driven infrastructure called Jilin Artificial-intelligence to col-
aided Materials-design Integrated Package (JAMIP). The organization of JAMIP abides by
the data lifecycle in computational materials informatics, from data generation to collec-
tion and learning, as shown in Figure 8. It provides tools for materials production, high-
throughput calculations, data extraction and management, and ML-based data mining.
Materials 2023, 16, 5977 12 of 30
Materials 2023, 16, x FOR PEER REVIEW 12 of 30

lection and learning, as shown in Figure 8. It provides tools for materials production,
high-throughput calculations, data extraction and management, and ML-based data min-
The authors demonstrated the usefulness of JAMIP in exploring materials informatics in
ing. The authors demonstrated the usefulness of JAMIP in exploring materials informatics
optoelectronic semiconductors, specifically halide perovskites. Hu et al. [75] proposed and
in optoelectronic semiconductors, specifically halide perovskites. Hu et al. [75] proposed
developed [Link] (accessed on 19 August 2023), a web-based materials infor-
and developed [Link] (accessed on 19 August 2023), a web-based materials
matics toolbox. The MaterialsAtlas platform includes tools for chemical validity check,
informatics toolbox. The MaterialsAtlas platform includes tools for chemical validity check,
formation energy and e-above-hull energy check, property prediction, screening of hypo-
formation energy and e-above-hull energy check, property prediction, screening of hypo-
theticalmaterials,
thetical materials,and
andutility
utility tools.
tools. The
The toolbox
toolbox lowers
lowers thethe barrier
barrier for materials
for materials scientists
scientists in
in data-driven
data-driven exploratory
exploratory materials
materials discovery.
discovery.

[Link]
Figure Overviewofofthe
theJAMIP
JAMIPcode
codeframework.
[Link]
programcomprises
comprises three
three major
major parts
parts based
based
on the material data’s lifecycle: data generation (blue), data collection (yellow), and data learning
on the material data’s lifecycle: data generation (blue), data collection (yellow), and data learning
(green).Reproduced
(green). Reproducedwith
withpermission
permissionfrom
from[74].
[74].

4.
4. ML
ML ininMaterials
MaterialsScience
Science
4.1.
4.1. Prediction of MaterialProperties
Prediction of Material Properties
ML
MLhas
hasgained
gainedprominence
prominenceininrecent
recentyears
yearsininpredicting
predictingmaterial
material properties
propertiesdue to to
due
its advantages of high generalization ability and fast computational speed. It
its advantages of high generalization ability and fast computational speed. It has beenhas been
successfully applied to predict the structure, adsorption, electrical, catalytic, energy storage,
successfully applied to predict the structure, adsorption, electrical, catalytic, energy stor-
and thermodynamic properties of materials. The prediction results could even reach the
age, and thermodynamic properties of materials. The prediction results could even reach
same accuracy as high-fidelity models with low computational costs.
the same accuracy as high-fidelity models with low computational costs.
4.1.1. Molecular Properties
4.1.1. Molecular Properties
In the past, it was very time consuming to predict molecular properties based on high-
In the density
throughput past, it generalization
was very timecalculations.
consuming ML to predict molecular
allows fast properties
and accurate basedofon
prediction
high-throughput density generalization calculations. ML allows fast and
the structure or properties of molecules, compounds, and materials. In materials science, accurate predic-
tion of thefactors,
solubility structure
suchorasproperties
Hansen and of molecules,
Hildebrand compounds,
solubility, areandcritical
materials. In materials
parameters for
science, solubility
characterizing factors,properties
the physical such as Hansen andsubstances.
of various HildebrandKurotani
solubility, are
et al. critical
[76] parame-
successfully
ters for characterizing
developed the physical
a solubility prediction modelproperties of various
with a unique substances.
ML method, Kurotaniin-phase
the so-called et al. [76]
successfully
DNN (ip-DNN).developed a solubility
This algorithm prediction
started with themodel
analysiswith
of ainput
unique
dataML method,NMR
(including the so-
called in-phase
information, DNN index,
refractive (ip-DNN). This algorithm
and density). startedwas
The solubility withthen
thespeculated
analysis ofininput data
a multi-
step approach
(including NMR by information,
predicting intermediate elements,
refractive index, such as molecular
and density). components
The solubility was thenand spec-
molecular
ulated in adescriptors.
multi-stepAn intermediate
approach regression
by predicting model was also
intermediate utilized
elements, to improve
such the
as molecular
accuracy
components of the
andprediction.
molecularAdescriptors.
website dedicated to the established
An intermediate regressionsolubility
model was prediction
also uti-
methods has also the
lized to improve been developed,
accuracy of thewhich is available
prediction. free ofdedicated
A website charge. Liang
to the et al. [77]
established
proposed a generalized ML method based on ANNs to predict polymer compatibility
solubility prediction methods has also been developed, which is available free of charge. (the
total miscibility of polymers with each other at the molecular scale).
Liang et al. [77] proposed a generalized ML method based on ANNs to predict polymer The authors built a
database by collecting
compatibility (the totaldata from scattered
miscibility literature
of polymers withthrough natural
each other at thelanguage
molecular processing
scale). The
techniques. By using the proposed method, predictions could be made
authors built a database by collecting data from scattered literature through natural based on the basiclan-
guage processing techniques. By using the proposed method, predictions could be made
based on the basic molecular structure of the blended polymers and the blended
Materials 2023, 16, 5977 13 of 30

molecular structure of the blended polymers and the blended compositions (as an auxiliary).
This generalized approach yielded some results in illustrating polymer compatibility. A
prediction accuracy of no less than 75% was achieved on a dataset containing 1400 entries
in their model. Zeng et al. [78] developed an atomic table CNN that could predict the band
gap and ground energy. The model accuracy exceeded that of standard DFT calculations.
Furthermore, this model could accurately predict superconducting transition temperatures
and distinguish between superconductors and non-superconductors. With the help of
this model, 20 potential superconductor compounds with high superconducting transition
temperatures were screened out.

4.1.2. Band Gap


The band gap size not only determines the energy band structure of a material but
also affects its electronic structure and optical properties. Recently, researchers have
applied ML to forecast the band gap of various materials. Venkatraman [79] developed an
algorithm for band gap prediction based on a rule-based ML framework. With descriptors
derived from elemental compositions, this model accurately and quickly predicted the
band gap of various materials. After testing on two independent sets, this model obtained
squared correlations > 0.85, with errors smaller than those of most density generalization
calculations, improving the material screening performance. Xu et al. [80] developed an
ML model called support vector regression (SVR) for predicting the band gaps of polymers.
They used training data obtained from DFT computations and generated descriptors
using Dragon software. After feature selection, the SVR model using 16 key features
achieved high accuracy in predicting polymer band gaps. The SVR model with a Gaussian
kernel function performed the best, with a determination coefficient (R2 ) of 0.824 and a
root mean square error (RMSE) of 0.485 in leave-one-out cross-validation. The authors
also provided correlation analysis and sensitivity analysis to understand the relationship
between the selected features and the band gaps of polymers. Several polymer samples
with targeted band gaps were designed based on the analysis and validated through DFT
calculations and model predictions. Espinosa et al. [81] proposed a vision-based system
to predict the electronic band gaps of organic molecules using deep learning techniques.
The system employed a multichannel 2D CNN and a 3D CNN to recognize and classify
2D projected images of molecular structures. The training and testing datasets used in the
research were derived from the Organic Materials Database (OMDB-GAP1). The results
showed that the proposed CNN model achieved a mean absolute error of 0.6780 eV and
an RMSE of 0.7673 eV, outperforming other ML methods based on conventional DFT.
These findings demonstrate the potential of CNN models in materials science applications
using orthogonal image projections of molecules. Wang et al. [82] explored the use of ML
techniques to accurately predict the band gaps of semiconductor materials. The authors
applied a stacking approach, which combined the outputs of multiple baseline models, to
enhance the performance of band gap regression. The effectiveness of different models
was tested using a benchmark dataset and a newly established complex database. The
results showed that the stacking model had the highest R2 value in both datasets, indicating
its superior performance. The improvement percentages of various evaluation metrics
for the stacking model compared to other baseline models range from 3.06% to 33.33%.
Overall, the research demonstrated the excellent performance of the stacking approach in
band gap regression. On the basis of generalized gradient approximation (GGA) band gap
information of crystal structures and materials, Na et al. [83] established an ML method
that used the tupleswise graph neural network (TGNN) algorithm for the accurate band
gap prediction of crystalline compounds. The TGNN algorithm showed strong superiority
in predicting the band gap of four different open databases. It has better accuracy for
48,835 samples of G0 W0 (a widely used technique in which the self-energy is expressed
as the convolution of a noninteracting Green’s function (G0 ) and a screened Coulomb
interaction (W0 ) in the frequency domain) band gaps than the standard density generalized
Materials 2023, 16, 5977 14 of 30

Materials 2023, 16, x FOR PEER REVIEW 14 of 30


theory without high computational costs. Moreover, this model could be extended to
project other valuable properties.
4.1.3. Energy Storage Performance
4.1.3. Energy Storage Performance
Energy storage is a key step in determining the efficiency, stability, and reliability of
Energy storage is a key step in determining the efficiency, stability, and reliability
power supply systems [84]. Exploring the energy storage performance of materials is crit-
of power supply systems [84]. Exploring the energy storage performance of materials
ical to energy storage, and ML accelerates the exploration process. Feng et al. [85] collected
is critical to energy storage, and ML accelerates the exploration process. Feng et al. [85]
over one thousand composite energy storage performance data points from the open lit-
collected over one thousand composite energy storage performance data points from the
erature and utilized ML to analyze and build a predictive model. The prediction accura-
open literature and utilized ML to analyze and build a predictive model. The prediction
cies of the RF, SVM, and neural network were 84.1%, 80.9%, and 70.6%, respectively. They
accuracies of the RF, SVM, and neural network were 84.1%, 80.9%, and 70.6%, respectively.
then added processed visual information data of the composite into the dataset, resulting
They then added processed visual information data of the composite into the dataset,
in improved prediction accuracies of 91.9%, 68.9%, and 81.6% for the three models, re-
resulting in improved prediction accuracies of 91.9%, 68.9%, and 81.6% for the three
spectively. This demonstrated
models, respectively. that the dispersion
This demonstrated that theof the filler in
dispersion ofthe
thematrix
filler inis the
an important
matrix is
factor affecting the maximum energy storage density of the composite.
an important factor affecting the maximum energy storage density of the composite. The The authors also
analyzedalso
authors theanalyzed
weights of theeach descriptor
weights in the
of each RF model
descriptor andRF
in the explored
model the andeffects of vari-
explored the
ous parameters on the energy storage of the material. Figure 9 shows
effects of various parameters on the energy storage of the material. Figure 9 shows the the logic diagram of
their ML models. Yue et al. [86] utilized the packing dielectric constant,
logic diagram of their ML models. Yue et al. [86] utilized the packing dielectric constant, packing size, and
packing size,
packing contentandaspacking
descriptors to predict
content the energy
as descriptors storagethe
to predict density
energy of storage
polymerdensity
matrix
composites. High-throughput random breakdown simulations
of polymer matrix composites. High-throughput random breakdown simulations were were performed on 504
datasets. The
performed simulation
on 504 [Link] were thenresults
The simulation applied as an
were MLapplied
then databaseas anandML combined
databasewith
and
classical dielectric prediction equations. They experimentally validated
combined with classical dielectric prediction equations. They experimentally validated the the predictions,
including the
predictions, dielectric
including theconstant
dielectricand breakdown
constant strength. strength.
and breakdown This work provides
This insights
work provides
into the design and fabrication of polymer matrix composites with
insights into the design and fabrication of polymer matrix composites with enhanced enhanced energy den-
sity for applications in capacitive energy storage. Ojin et al. [87] built
energy density for applications in capacitive energy storage. Ojin et al. [87] built four four traditional ML
models andML
traditional two graph and
models neural
twonetwork models.
graph neural Through
network them,Through
models. 32,026 heatthem,capacity
32,026struc-
heat
tures were
capacity predicted
structures using
were a high-precision
predicted deep graph attention
using a high-precision deep graphnetwork. Additionally,
attention network.
the correlation
Additionally, thebetween heatbetween
correlation capacityheatandcapacity
structure
and descriptors was inspected.
structure descriptors A total of
was inspected.
22 structures were predicted to have high heat capacity, and the results
A total of 22 structures were predicted to have high heat capacity, and the results were were further vali-
dated by DFT analysis. Through the combination of ML and
further validated by DFT analysis. Through the combination of ML and minimal DFT minimal DFT queries, this
study provides
queries, this studya path to accelerating
provides a path to the discoverythe
accelerating of new thermal
discovery of energy
new thermalstorageenergy
mate-
rials. materials.
storage

Figure 9.
Figure 9. Logic
Logic diagram
diagram of
of predicting
predicting the
the maximum
maximum energy
energy density
density and
and exploring
exploring the
the potential
potential
effective structure of composites through the ML method, reproduced with permission from [85].
effective structure of composites through the ML method, reproduced with permission from [85].

4.1.4.
4.1.4. Structural
Structural Health
Health
Structural
Structural health monitoring
health monitoring (SHM)
(SHM) utilizes
utilizes engineering,
engineering, scientific,
scientific, and
and foundational
foundational
knowledge
knowledge to prevent
prevent damage to property
property and
and life.
life. The core ofof the
the field
field of
of construction
construction
informatics
informatics isis the
the transmission,
transmission, processing,
processing, and
and visualization
visualization of architectural information,
of architectural information,
providing
providing effective
effective methods
methods forfor monitoring
monitoring structural
structural changes
changes [88,89]. ML provides
[88,89]. ML provides effec-
effec-
tive methods for monitoring structural changes. Dang et al. [90] proposed
tive methods for monitoring structural changes. Dang et al. [90] proposed a cloud-based a cloud-based
digital
digital twin
twin framework
framework for for SHM
SHM employing
employing deepdeep learning.
learning. The
The framework
framework consists
consists of
of
physical components, device measurements, and digital models formed
physical components, device measurements, and digital models formed by combining dif- by combining
different sub-models
ferent sub-models includingmathematical,
including mathematical,finite
finiteelement,
element,and
and MLML sub-models.
sub-models. The The data
data
interactions among the physical structure, digital model, and human interventions were
enhanced by using cloud computing infrastructure and a user-friendly web application.
Materials 2023, 16, 5977 15 of 30

interactions among the physical structure, digital model, and human interventions were
enhanced by using cloud computing infrastructure and a user-friendly web application.
The feasibility of the framework was demonstrated through case studies of the damage de-
tection of model bridges and real bridge structures utilizing deep learning algorithms, with
a high accuracy of 92%. Dong et al. [91] discussed the use of the eXtreme gradient boosting
(XGBoost) algorithm for predicting concrete electrical resistivity in SHM (Figure 10a). The
proposed XGBoost-algorithm-based prediction model considers all potential influencing
factors simultaneously. A database of 800 experimental instances was used to train and test
the model. The results showed that the XGBoost model achieved satisfactory predictive
performance. The study also identified the importance of curing age and cement content
in electrical resistivity measurement results. The XGBoost algorithm was chosen for its
high performance, ease of use, and better prediction accuracy than other algorithms. The
bond effect between the reinforcement and concrete guarantees the combined action of the
two materials. This is a critical factor that affects the mechanical properties of reinforced
concrete components and structures, e.g., bearing capacity and ductility [92]. Gao et al. [93]
developed a new solution for evaluating the bond strength of an FRP using AI-based
models. Two hybrid models, the imperialist competitive algorithm (ICA)-ANN and the
artificial bee colony (ABC)-ANN, were designed and compared. The results showed that
the ICA-ANN model had a higher predictive ability than the ABC-ANN model. The pro-
posed hybrid models can be used as a suitable substitute for empirical models in evaluating
FRP bond strength in concrete samples. Li et al. [94] utilized ML approaches to estimate
the bond strength between ultra-high-performance concrete (UHPC) and reinforcing bars.
A new database was created by integrating data from multiple published works. Nine
ML models, including linear models, tree models, and ANNs, were implemented to train
bond strength estimators based on the database. The results showed that the ANN and
RF models achieved the highest estimation performances, surpassing empirical formulas.
The study also analyzed the relative importance of different factors in determining bond
strength. Overall, the research provides a data-driven approach to estimating bond strength
and contributes to the understanding of bond performance between UHPC and reinforcing
bars. Su et al. [95] applied three ML approaches (multiple linear regression, SVM, and
ANN) to predict the interfacial bond strength between FRPs and concrete (Figure 10b). They
trained these models using two datasets containing experimental results from single-lap
shear tests, employed random search and grid search to find the optimal hyperparameters,
and analyzed input variables’ contributions using partial dependence plots. They also
developed a stacking strategy to improve prediction accuracy. The results showed that the
SVM approach had the best accuracy and efficiency. They concluded that ML methods
are feasible and efficient for predicting the bond strength of FRP laminates in reinforced
concrete structures.

4.1.5. Nanomaterial Toxicity


It has been proven that ML can be used to identify nanomaterial properties and expo-
sure conditions that influence cellular and organism toxicity, thus providing information
required for risk assessment and safe-by-design approaches in the development of new
nanomaterials [96]. Huang et al. [97] combined ML with high-throughput in vitro bioassays
to develop a model to predict the toxicity of metal oxide nanoparticles to immune cells, as
shown in Figure 11. In the training, test, and experimental validation sets, the ML model
displayed prediction accuracies of 97%, 96%, and 91%, respectively. ML methods were
used to identify features that encode information on immune toxicity. These features are
crucial for the scientific design of future experiments and for the accurate depiction of
nanotoxicity. According to Gousiadoua et al. [98], advanced ML techniques were applied
to create nano quantitative structure–activity relationship (QSAR) tools for modeling the
toxicity of metallic and metal oxide nanomaterials, both coated and uncoated, with various
core compositions tested on embryonic zebrafish at various dosage concentrations. Based
on both computed and experimental descriptors, the scientists identified a set of properties
Materials 2023, 16, 5977 16 of 30

most relevant for assessing nanomaterial toxicity and successfully correlated these prop-
erties with zebrafish physiological responses. It has been concluded that for the group of
metal and metal oxide nanomaterials, the core chemical composition, concentration, and
properties are influenced by the nanomaterial surface and medium composition (such as
zeta potential and agglomerate size), which have a significant impact on toxicity, even
though the ranking of different variables is subject to variation in the analytical method
and data model. Generalized nano-QSAR ensemble models offer a promising framework
for predicting the toxicity potential of new nanomaterials. Liu et al. [99] presented a meta-
analysis of phytosynthesized silver nanoparticles (AgNPs) with heterogeneous features
using DTs and RFs. The researchers found that exposure regime (including the time and
dose), plant family, and cell type were the most important predictors for cell viability for
green AgNPs. In addition, a discussion of the potential effects of major variables (cell
assays, inherent nanoparticle properties, and reaction parameters used in biosynthesis)
on AgNP-mediated cytotoxicity and model performance was presented to provide a basis
Materials 2023, 16, x FOR PEER REVIEW
for future research. The findings of this study may assist future studies in improving the 16 of 3
design of experiments and the development of virtual models or optimizations of green
AgNPs for specific applications.

Figure10.
Figure 10. (a)
(a) Schematic
SchematicofofXGBoost
XGBoost trees,
trees, adapted
adapted withwith permission
permission fromfrom
[91].[91]. (b) ML
(b) ML model con
model
struction process, adapted with permission from [95].
construction process, adapted with permission from [95].

4.1.5. Nanomaterial Toxicity


It has been proven that ML can be used to identify nanomaterial properties and ex
posure conditions that influence cellular and organism toxicity, thus providing infor
mation required for risk assessment and safe-by-design approaches in the developmen
of new nanomaterials [96]. Huang et al. [97] combined ML with high-throughput in vitro
bioassays to develop a model to predict the toxicity of metal oxide nanoparticles to im
for cell viability for green AgNPs. In addition, a discussion of the potential effects
variables (cell assays, inherent nanoparticle properties, and reaction parameters
biosynthesis) on AgNP-mediated cytotoxicity and model performance was pres
provide a basis for future research. The findings of this study may assist future s
Materials 2023, 16, 5977 improving the design of experiments and the development of virtual17modelsof 30 or o
tions of green AgNPs for specific applications.

Figure
Figure 11.11. Schematic
Schematic workflow
workflow of data compilation,
of data compilation, descriptor
descriptor generation, generation,
machine learningmachine
model- learn
ing, experimental
eling, validation,
experimental and mechanism
validation, interpretation,interpretation,
and mechanism reproduced with permission
reproduced fromwith
[97]. permiss
[97]. Adsorption Performance of Nanomaterials
4.1.6.
Because of their high surface area, ease of functionalization, and affinity toward a
4.1.6.
wide Adsorption
range Performance
of pollutants, nanomaterialsofare Nanomaterials
excellent adsorbents [100]. Moosavi et al. [101]
appliedBecause
four machine learning methods
of their high surface area, to model ease
dye adsorption on 16 activatedand
of functionalization, carbonaffinity t
adsorbents and determined the relationship between adsorption capacity and activated
wide range of pollutants, nanomaterials are excellent adsorbents [100]. Moosavi et
carbon parameters. The results indicated that agro-waste characteristics (pore volume,
applied
surface four
area, pH,machine
and particle learning methods
size) contributed to model
50.7% dye adsorption
to the adsorption [Link] 16 activated
Among
adsorbents
the agro-wasteand determined
characteristics, porethe relationship
volume and surface between
area were adsorption capacity and a
the most important
influencing variables, while
carbon parameters. Theparticle
resultssizeindicated
had a limited
thatimpact. With a hypothetical
agro-waste set of (pore
characteristics
approximately
surface area, pH, and particle size) contributed 50.7% to the adsorptionand
130,000 structures of metal–organic frameworks (MOFs) with methane efficiency
carbon dioxide adsorption data at different pressures, Guo et al. [102] established models
the agro-waste characteristics, pore volume and surface area were the most impo
for estimating gas adsorption capacities using two deep learning algorithms, multilayer
fluencing(MLPs)
perceptrons variables, while
and long particle
short-term size had
memory (LSTM) a limited
[Link].
The modelsWith
were a eval-
hypothetic
approximately
uated by performing130,000 structures
ten iterations of 10-foldofcross-validations
metal–organic andframeworks (MOFs) with
100 holdout validations.
The
and carbon dioxide adsorption data at different pressures, Guo etpredic-
performance of the MLP and LSTM models was similar with high accuracy of al. [102] est
tion. Those models that predicted gas adsorption at a higher pressure performed better
models for estimating gas adsorption capacities using two deep learning algorithm
than those that predicted gas adsorption at a lower pressure. In particular, deep learning
tilayerwere
models perceptrons
more accurate(MLPs) andmodels
than RF long reported
short-term memory
in the literature(LSTM) networks. The
when predicting
were
gas evaluated
adsorption by performing
capacities ten iterations
at low pressures. of 10-fold
Deep learning cross-validations
algorithms were found to beand 100
highly
validations. The performance of the MLP and LSTM models gas
effective in generating models capable of accurately predicting the wasadsorption
similar with hi
capacities of MOFs.
racy of prediction. Those models that predicted gas adsorption at a higher press
formed
4.2. better
Accelerated than those
Materials that
Synthesis andpredicted
Design gas adsorption at a lower pressure. In pa
deep learning models were more accurate than
In addition to being widely utilized for predicting RF models
material reported
properties, inplays
ML also the literatu
apredicting gas
pivotal role in theadsorption capacities
synthesis of new [Link]
low pressures. DeepML
the past few years, learning
has madealgorithm
Materials 2023, 16, 5977 18 of 30

significant progress in the exploration of novel materials, such as highly efficient molecular
organic light-emitting diodes [103], low thermal hysteresis shape memory alloys [104],
and piezoelectric materials with large electrical strain [105]. The use of ML for materials
synthesis not only significantly speeds up novel material discovery but also provides
insight into the basic composition changes in materials from big data.

4.2.1. Chalcogenide Materials


Chalcogenide materials can be used in a variety of photovoltaic and energy devices,
including light-emitting diodes, photodetectors, and batteries. ML has promoted the
development of high-performance chalcogenide materials [106]. Li et al. [107] proposed an
ML model based on an RF algorithm for speculating the formation of ABX3 and A2 B0 B00 X6
compound chalcogenides. With geometric and electrical parameters, the RF classification
model reached 96.55% accuracy for ABX3 samples and 91.83% accuracy for A2 B0 B00 X6
samples. A total of 241 ABX3 chalcogenides with a 95% probability of formation were
filtered from 15,999 candidate compounds, and a total of 1131 A2 B0 B00 X6 chalcogenides
with a 99% probability of formation were filtered from 417,835 candidate compounds. The
method presented in their work could offer valuable enlightenment for the acceleration of
discovering perovskites. Liu et al. [108] used data from 397 ABO3 compounds and nine
parameters (e.g., tolerance factor and octahedral factor) as input variables for ML. The
gradient-enhanced DT obtained by training was compared as the optimal model by 10-fold
cross-validation of the average accuracy. A total of 331 chalcogenides were filtered by the
model from 891 data points with a classification accuracy of 94.6%. Omprakash et al. [109]
compiled a model including organometallic salt chalcogenides to 2D chalcocite and its
corresponding band gaps. An ML model for predicting all types of chalcocite band gaps
was then trained using a graphical representation learning technique. The model could
accurately estimate the band gap within a few milliseconds with an average absolute
error of 0.28 eV. Wang et al. [110] applied unsupervised learning to discover quaternary
chalcogenide semiconductors (I2 -II-IV-X4 ) and were successful in screening eight of these
materials with good photoconversion efficiency despite a data shortage. This method
shortens the material screening cycle and facilitates rapid material discovery.

4.2.2. Catalytic Materials


In traditional experiments, it is difficult to design efficient catalytic materials in a
short time because a clear reaction mechanism is required [111]. ML can rapidly extract the
relationship between the structure and performance of catalytic materials and effectively
expedite the development process of new catalytic materials. Zhang et al. [112] employed
a gradient boosting algorithm to build an ML model. The model utilized four key stability
and catalytic features of graphene-loaded single-atom catalysts as targets to find catalytic
materials suitable for electro-hydrogenation nitrogen reactions. With this model, a total of
45 catalytic materials with efficient catalytic performance were successfully screened from
1626 samples. The model could be operated for the rapid screening of other electrocata-
lysts. Figure 12 illustrates their computational framework. Wei et al. [113] developed an
ML model, which was applied in a Bayesian optimization framework to obtain molyb-
denum disulfide (MoS2 ) catalysts with stable hydrogen reaction activity. To explore the
structure–property relationship of the samples optimized by the ML technique, nine elec-
trochemical characterizations were performed to verify the results, including SEM, TEM,
XRD, and XPS. A strong correlation was found between the structure of the optimized
MoS2 and its hydrogen evolution reaction performance. Hueffel et al. [114] reported an
unsupervised ML workflow that uses only five experimental data points, which could be
used to accelerate the recognition of binuclear palladium (Pd) catalysts. Based on their
method, some phosphine ligands were successfully predicted and experimentally verified
from 348 ligands, including those that had never been synthesized before, which formed
binuclear Pd(I) complexes on Pd(0) and Pd(II) species. Their strategy plays an important
optimized MoS2 and its hydrogen evolution reaction performance. Hueffel et al. [114] re
ported an unsupervised ML workflow that uses only five experimental data points, which
could be used to accelerate the recognition of binuclear palladium (Pd) catalysts. Based
Materials 2023, 16, 5977 on their method, some phosphine ligands were successfully predicted and experimentally19 of 30
verified from 348 ligands, including those that had never been synthesized before, which
formed binuclear Pd(I) complexes on Pd(0) and Pd(II) species. Their strategy plays an im
portant
role role in the
in studying studying themechanisms
formation formation mechanisms
of Pd catalystof Pd catalyst
species, as wellspecies, as well as th
as the further
further integration
integration of ML intoofcatalytic
ML into catalytic research.
research.

Figure12.
Figure 12. Catalyst
Catalyststructures,
structures,target
target properties,
properties, andand computational
computational framework.
framework. (a) Structural rep
(a) Structural
resentation ofofthree-coordinated
representation three-coordinated andand four-coordinated
four-coordinated configurations.
configurations. Letter
Letter “M” “M” represents
represents the th
centralmetal
central metalatom,
atom,andand letter
letter “C”“C” represents
represents the coordinating
the coordinating atom ofatom ofTarget
M. (b) M. (b)properties
Target properties
for fo
describingthe
describing the 2 fixation
N2Nfixation performance
performance ofcatalyst.
of the the catalyst.
(c) ML(c) ML screening
screening and descriptor
and descriptor building buildin
frameworkofof
framework their
their work.
work. Reproduced
Reproduced withwith permission
permission from [112].
from [112].

4.2.3. Superconducting Materials


4.2.3. Superconducting Materials
Superconductivity, intrinsically regulated by finite phonon-coupled electron–electron
Superconductivity,
attractions, has aroused decades intrinsically
of intenseregulated by finite
research interest phonon-coupled
in condensed electron–elec
matter physics.
The development and prediction of upcoming superconducting materials with high critical matte
tron attractions, has aroused decades of intense research interest in condensed
physics. Theare
temperatures development andapplications.
essential in many prediction of upcoming
ML-guided superconducting
iterative experimentation materials
may with
high critical
outperform temperatures
standard are essential
high-throughput in many
screening applications.
for discovering ML-guided
breakthrough iterative
materials in experi
high-temperature superconductors
mentation may outperform [115,116].
standard Zhang et al. [117]
high-throughput developed
screening forandiscovering
integrated break
ML model to accurately and robustly predict the critical
through materials in high-temperature superconductors [115,116]. temperature (Tc ) of superconduct-
Zhang et al. [117] de
ing materials (Figure 13a). They used open-source materials data, ML models, and data
veloped an integrated ML model to accurately and robustly predict the critical tempera
mining methods to explore the correlation between chemical features and Tc values. The
ture (Tc) of superconducting materials (Figure 13a). They used open-source materials data
integrated model combined three basic algorithms (gradient boosting decision tree, extra
ML and
tree, models, and databoosting
light gradient miningmachine)
methods to to explore
improve thethe correlation
prediction between
accuracy. chemical fea
The model
tures and T
achieved an R of 95.9% and an RMSE of 6.3 K. The study also identified the importance of(gradien
2
c values. The integrated model combined three basic algorithms
boosting
various decision
material tree,inextra
features tree, and
Tc prediction, light
with gradient
thermal boosting
conductivity machine)
playing to improve
a critical role. th
prediction
The integratedaccuracy.
model was Theused
model achieved
to screen out an R2 of 95.9%
potential and an RMSE
superconducting of 6.3 with
materials K. The study
Talso
c values beyondthe
identified 50.0importance
K. This research provides
of various insightsfeatures
material for accelerating the exploration
in Tc prediction, with therma
of high-Tc superconductors.
conductivity playing a criticalRoter et al.
role. [118]
The used MLmodel
integrated to predict
wasnew
used superconductors
to screen out potentia
and their critical temperatures. They constructed a database of superconductors and their
superconducting materials with Tc values beyond 50.0 K. This research provides insight
chemical compositions and applied this information to train ML models. They achieved
for accelerating the exploration of high-Tc superconductors. Roter et al. [118] used ML to
an R2 of approximately 0.93, which was comparable to or higher than similar estimates
based onnew
predict othersuperconductors
AI techniques. They and also
theirdiscussed
critical temperatures. They
factors that limit constructed
learning and sug-a databas
of superconductors
gested possible ways and their chemical
to overcome them. Thecompositions
researchersandusedapplied this information
both unsupervised and to train
ML models.
supervised ML They achieved
techniques, an Rsingular
including 2 of approximately 0.93, which was comparable to o
value decomposition and KNN, to improve
higher
their than similar
models’ [Link] based aonclassification
They achieved other AI techniques.
accuracy ofThey
96.5%also an R2 of factor
anddiscussed
approximately 0.93 for
that limit learning andpredicting
suggestedcritical temperatures.
possible ways toThey also employed
overcome them. The their models
researchers used
to predict several new superconductors with high critical temperatures. However, the
both unsupervised and supervised ML techniques, including singular value decomposi-
tion and KNN, to improve their models’ accuracy. They achieved a classification accuracy
Materials 2023, 16, 5977 of 96.5% and an R2 of approximately 0.93 for predicting critical temperatures.20They of 30 also
employed their models to predict several new superconductors with high critical temper-
atures. However, the authors noted that incorrect entries in the database can lead to out-
liers in noted
authors the predictions. Pereti
that incorrect et [Link][119]
entries proposed
the database anlead
can MLtoapproach to the
outliers in identify new super-
predictions.
conducting
Pereti materials.
et al. [119] They
proposed an utilized
ML approach DeepSet technology,
to identify new which allows them
superconducting to input the
materials.
chemical
They utilizedconstituents of the compounds
DeepSet technology, which allows without
them topredetermined ordering
input the chemical (Figure
constituents of 13b).
the
Thecompounds
method was without predetermined
successful ordering
in classifying (Figure as
materials 13b). The method wasand
superconducting successful
quantifying
intheir
classifying
critical materials as superconducting
temperature. The trained neural andnetwork
quantifying
was their critical
then used totemperature.
search through a
The
mineralogical database for candidates that might be [Link]
trained neural network was then used to search through a mineralogical for
Three materials
candidates
were selectedthat might be superconducting.
for experimental Three materials
characterization, were selected for experimental
and superconductivity was confirmed in
characterization, and superconductivity was confirmed in two of them. This was the first
two of them. This was the first time a superconducting material was identified using AI
time a superconducting material was identified using AI methods. The results demon-
methods. The results demonstrated the effectiveness of the DeepSet network in predicting
strated the effectiveness of the DeepSet network in predicting the critical temperatures of
the critical temperatures of superconducting materials.
superconducting materials.

Figure 13. (a) Workflow of the integrated model-based ML methods for accurate Tc prediction and
new superconductor material mining, adapted with permission from [117]. (b) A schematic layout of
the DeepSet architecture, adapted with permission from [119].
Figure 13. (a) Workflow of the integrated model-based ML methods for accurate Tc prediction and
new superconductor material mining, adapted with permission from [117]. (b) A schematic layout
Materials 2023, 16, 5977 of the DeepSet architecture, adapted with permission from [119]. 21 of 30

4.2.4. Nanomaterial Outcome Prediction


4.2.4. Rapid advancements
Nanomaterial Outcome in Prediction
materials synthesis techniques have led to more and more
attention being paid to nanomaterials,
Rapid advancements in materials synthesis including nanocrystals,
techniques have led nanorods,
to morenanoplates,
and more
nanoclusters,
attention beingand paidnanocrystalline
to nanomaterials, thinincluding
films. Materials of this nanorods,
nanocrystals, class offer nanoplates,
enhanced physi- nan-
cal and chemical
oclusters, tunability across
and nanocrystalline a range
thin films. of systems,
Materials of this including
class offer inorganic semiconduc-
enhanced physical and
tors, metals,
chemical and molecular
tunability crystals.
across a range of A nanomaterial
systems, including is defined
inorganic as asemiconductors,
material with a metals,
dimen-
sionmolecular
and smaller than 100 nanometers
crystals. A nanomaterial in at least one dimension.
is defined as a material Unlike
withbulk materials,
a dimension nano-
smaller
materials possess different physical and chemical properties due
than 100 nanometers in at least one dimension. Unlike bulk materials, nanomaterials possess to their unique size and
shape. This
different technology
physical has a broad
and chemical array of
properties dueapplication prospects,
to their unique including
size and [Link]-
tech-
sion and storage of energy, the restoration of water, medical treatment,
nology has a broad array of application prospects, including the conversion and storage of and the storage
and processing
energy, of data.
the restoration of water, medical treatment, and the storage and processing of data.
Using experimental
Using experimental data, data, Xie
Xie etet al.
al. [120]
[120]reported
reportedthe thedevelopment
development of of an
anML-aided
ML-aided
method for predicting the crystallization tendency of metal–organic
method for predicting the crystallization tendency of metal–organic nanocapsules (MONCs). nanocapsules
A(MONCs).
prediction A accuracy
prediction ofaccuracy
>91% was of achieved
>91% wasby achieved
using the by XGBoost
using the model.
XGBoost model. Fur-
Furthermore,
thermore,
they they synthesized
synthesized a set of newacrystalline
set of newMONCs crystalline
using MONCs usingfeatures
the derived the derived features
and chemical
and chemical
hypotheses fromhypotheses
the XGBoost frommodel.
the XGBoost
The resultsmodel. The study
of this resultsdemonstrate
of this studythat demonstrate
ML algo-
that ML
rithms algorithms
can can assist
assist chemists chemists
in finding the in finding
optimal the optimal
reaction reaction
parameters fromparameters from a
a large number
large
of number ofparameters
experimental experimental more parameters
efficiently. more
Figureefficiently.
14 shows aFigureschematic14 shows a schematic
representation of
representation
the working flow. of the working
Pellegrino et flow.
al. [121]Pellegrino
tuned the et TiO
al. [121] tuned themorphology
2 nanoparticle TiO2 nanoparticle
using
morphology using
hydrothermal hydrothermal
treatment. treatment.
In their work, In their work,
an experimental an was
design experimental
employeddesign was
to investi-
employed
gate to investigate
the influence the influence
of relevant of relevant process
process parameters parameters
on the synthesis on the synthesis
outcome, enabling out-ML
methods to develop
come, enabling predictive
ML methods to models. After validation
develop predictive models. and training,
After the models
validation were
and training,
capable of accurately predicting the synthesis outcome in terms
the models were capable of accurately predicting the synthesis outcome in terms of nano-of nanoparticle size, poly-
dispersity, and aspect ratio. They presented a synthesis method that
particle size, polydispersity, and aspect ratio. They presented a synthesis method that al- allows the continuous
and
lowsprecise control ofand
the continuous nanoparticle morphology.
precise control This method
of nanoparticle affords the
morphology. possibility
This method to tune
affords
the
theaspect ratioto
possibility over
tunea large rangeratio
the aspect fromover1.4 (perfect truncated
a large range from bipyramids) to 6 (elongated
1.4 (perfect truncated bipyr-
nanoparticles) and a length
amids) to 6 (elongated from 20 to and
nanoparticles) 140 nm.a length from 20 to 140 nm.

Figure14.
Figure [Link]
Schematicrepresentation
representationofofthe
theworking
workingflow
flowwhen
when machine
machine learning
learning models
models areare incor-
incorpo-
porated into the prediction of the crystallization propensity of MONCs, with permission from
rated into the prediction of the crystallization propensity of MONCs, with permission from [120]. [120].

4.2.5. Nanomaterial Synthesis


Nanomaterial synthesis often involves multiple reagents and interdependent exper-
imental conditions. Each experimental variable’s contribution to the final product is
generally determined through trial and error, along with intuition and experience. The
Materials 2023, 16, 5977 22 of 30

process of identifying the most efficient recipe and reaction conditions is therefore time
consuming, laborious, and resource intensive [122]. In a recent study, Erick et al. [123] used
SVM classification and regression models to predict the synthesis of CsPbBr3 nanosheets
with controlled layer thicknesses. The SVM classification is shown to accurately predict the
likelihood that CsPbBr3 synthesis would form a majority population of quantum-confined
nanoplatelets. Additionally, SVM regression can be used to determine the average thickness
of the synthesis of CsPbBr3 nanoplatelets with sub-monolayer accuracy. Epps et al. [124]
proposed a method that is based on ML experiment selection and high-efficiency au-
tonomous flow chemistry. The approach utilized SVM regression to predict the thickness of
the nanoplatelets and was shown to be accurate and reliable. Using this method, inorganic
perovskite quantum dots (QDs) in flow were synthesized autonomously. By using less than
210 mL of starting solutions and without user selection, this method synthesized precision
tailored QD compositions within 30 h. This would enable the commercialization of these
QDs, as well as their integration into various applications. Furthermore, the method could
be used for other types of nanomaterials, such as nanorods and nanowires.

4.2.6. Inverse Design of Nanomaterials


As opposed to the direct approach that leads from the chemical space to the desired
properties, inverse design starts with desired properties as the “input” and ends with
chemical space as the “output” [125]. In the field of nanomaterials, the complexity of
inverse design is enhanced by the finite dimensions and variety of shapes, resulting in
a larger design space [126]. The inverse design of nanomaterials was quite challenging
in the past. The inverse design of nanomaterials could be explored using interpretable
relationships between structure and property generated by ML methods. A new inverse
design method for metal nanoparticles based on deep learning was proposed and demon-
strated by Wang et al. [127]. In comparison to the least squares method, the calculated
results indicated that the inverse design method utilizing the back-propagation network
had greater adaptability, a smaller minimum error, and can be adjustable based on S pa-
rameters. Inverse design systems based on deep learning neural networks may be applied
to the inverse design of nanoparticles of different shapes. In another study, Li et al. [126]
demonstrated a novel approach to inverse design using multi-target regression methods
using RFs. A multi-target regression model was used with a precursory forward structure–
property prediction to capture the most important characteristics of a single nanoparticle
before the problem was inverted and a number of structural features were simultaneously
predicted. A general workflow has been demonstrated on two nanoparticle datasets, and
it has the capacity to predict rapid relationships between properties and structures for
guiding further research and development without the need for additional optimization
or high-throughput sampling. He et al. [128] employed a DNN to establish mappings
between the far-field spectra/near-field distribution and dimensional parameters of three
different types of plasmonic nanoparticles, including nanospheres, nanorods, and dimers.
Through the DNN, both the forward prediction of far-field optical properties and the
inverse prediction of nanoparticle dimensional parameters can be accomplished accurately
and efficiently. Figure 15 shows the structure of the reported machine learning model for
predicting optical properties and designing nanoparticles.
Materials 2023,
Materials 2023, 16,
16, 5977
x FOR PEER REVIEW 23
23 of 30
of 30

Figure 15.
Figure 15. Structures
Structures ofof machine
machine learning
learning models
models for
for predicting
predicting optical
optical properties
properties and
and designing
designing
nanoparticles. (a) Far- and near-field optical data obtained from the finite-difference time-domain
nanoparticles. (a) Far- and near-field optical data obtained from the finite-difference time-domain
(FDTD) simulations were used to train three different machine learning models: far-field spectra
(FDTD) simulations were used to train three different machine learning models: far-field spectra
and structural information for (i) structure classification, far-field spectra and dimensions for (ii) the
and structural information for (i) structure classification, far-field spectra and dimensions for (ii) the
spectral DNN, and near-field enhancement maps and dimensions for (iii) the E-field DNN. After
spectral
training,DNN,
machine and near-field
learning enhancement
models can be usedmaps and dimensions
to perform for (iii) the
forward prediction E-field
and/or [Link].
inverse After
training, machine learning models can be used to perform forward prediction and/or inverse
The solid and dashed red arrows represent the forward prediction and the inverse design process, design.
The solid and(b)
respectively. dashed red architecture
Detailed arrows represent
of thethe forward
three machineprediction
learningand the inverse
models design
in Figure process,
9a, with per-
mission from [128].
respectively. (b) Detailed architecture of the three machine learning models in Figure 9a, with
permission from [128].
5. Conclusions, Challenges, and Prospects
5. Conclusions, Challenges, and Prospects
This review discussed the use of machine learning (ML) in the field of materials sci-
ence This review discussed
for predicting materialthe use of machine
properties learning
and guiding (ML) insynthesis.
material the field ofThe
materials
reviewscience
briefly
for predicting material properties and guiding material synthesis. The
outlined the basic principles of ML and introduced commonly used algorithms and review briefly
their
outlined the basic
applications principles
in material of MLand
screening andproperty
introduced commonly
prediction. used
It also algorithms
presented theand their
research
applications in material screening and property prediction. It also presented the research
progress of ML in predicting material properties and guiding material synthesis. The
Materials 2023, 16, 5977 24 of 30

progress of ML in predicting material properties and guiding material synthesis. The review
suggested that ML can greatly reduce computational costs, shorten the development cycle,
and improve computational accuracy, making it a promising research approach in novel
materials screening and material property prediction.
It is important to note, however, that the following challenges still exist. Most ML
algorithms require large amounts of data to work properly. Even for the simplest problems,
thousands of examples are desired. Acquiring an effective dataset is critical for the research
and implementation of ML in materials science. However, data in materials science are
characterized by high acquisition costs, excessive concentration or dispersion, and a lack
of uniform processing standards. A dataset with a large amount of data, a uniform distri-
bution, and matching feature parameters is often extremely difficult to obtain. Although
material databases have greatly facilitated researchers’ access to data, many published data
have not been specified to date. The task of enriching existing databases is challenging. Text
mining techniques could be effective in rapidly collecting data scattered in the literature.
This approach could greatly enhance existing databases and create specialized databases.
The selection of features significantly affects the accuracy of ML models. Currently, the
use of manual feature engineering to filter features is often influenced by the researcher’s
experience and intuition. This approach may overlook some significant features. In contrast,
automated feature engineering automatically constructs new candidate features from the
data and selects the most appropriate features for model training, which could effectively
solve the current dilemma.
ML methods cannot replace traditional computational and experimental studies. Al-
though ML methods have shown remarkable promise in guiding the synthesis of novel
materials and predicting material properties, they are still mostly “black boxes” [108]. The
predicted results still need to be experimentally verified and the underlying physicochemi-
cal laws still need to be studied in depth. Therefore, ML can only perform some exploratory
tasks at present. With further improvement of theories and methods, however, ML might
eventually replace traditional experimental research by providing novel ideas and research
methods for the field of materials science. The application of ML in the field of materials
science and engineering is just the beginning, and its potential is endless in the future.

Author Contributions: Writing—original draft preparation, G.H. and Y.G.; writing—review and
editing, Z.N. and Y.C. All authors have read and agreed to the published version of the manuscript.
Funding: This work was supported by the Postgraduate Research and Practice Innovation Program of
Jiangsu Province (College Project), China, the Natural Science Foundation of Jiangsu Province, China
(Grant No. BK20200686), and the National Natural Science Foundation of China (Grant No. 52206257).
Conflicts of Interest: The authors declare that the research was conducted in the absence of any
commercial or financial relationships that could be construed as potential conflicts of interest.

Abbreviations
ABC Artificial bee colony
AFLOW Automatic Flow
AgNP Silver nanoparticle
AI Artificial intelligence
ANN Artificial neural network
CNN Convolutional neural network
COD Crystallography Open Database
CSD Cambridge Structural Database
DBM Deep Boltzmann machine
DBN Deep belief network
DFT Density functional theory
DNN Deep neural network
Materials 2023, 16, 5977 25 of 30

DT Decision tree
FDTD Finite-difference time-domain
FRP Fiber-reinforced polymer
GAN Generative adversarial network
GBRT Gradient boosted regression tree
GGA Generalized gradient approximation
GLOnet Global optimization network
HEA High-entropy alloy
HHV Higher heating value
ICA Imperialist competitive algorithm
ICSD Inorganic Crystal Structure Database
JAMIP Jilin Artificial-intelligence aided Materials-design Integrated Package
KNN K-nearest neighbor
LSTM Long short-term memory
MDN Mixture density network
ML Machine learning
MLP Multilayer perceptron
MOF Metal–organic framework
MONC Metal–organic nanocapsule
NMR Nuclear magnetic resonance
OMDB Organic Materials Database
OQMD Open Quantum Materials Database
QD Quantum dot
QSAR Quantitative structure–activity relationship
R-CNN Region-based CNN
ResNet Residual network
RF Random forest
RMSE Root mean square error
RNN Recurrent neural network
SEM Scanning electron microscope
SHM Structural health monitoring
SMILE Simplified molecular input line entry
SVM Support vector machine
SVR Support vector regression
TEM Transmission electron microscope
TGNN Tupleswise graph neural network
UHPC Ultra-high-performance concrete
XGBoost eXtreme gradient boosting
XPS X-ray photoelectron spectroscopy
XRD X-ray diffraction
µ-PTAMAM Micro-plasma transfer arc metal additive manufacturing

References
1. Lu, S.; Zhou, Q.; Ouyang, Y.; Guo, Y.; Li, Q.; Wang, J. Accelerated discovery of stable lead-free hybrid organic-inorganic
perovskites via machine learning. Nat. Commun. 2018, 9, 3405. [CrossRef] [PubMed]
2. Kolahalam, L.A.; Viswanath, I.K.; Diwakar, B.S.; Govindh, B.; Reddy, V.; Murthy, Y. Review on nanomaterials: Synthesis and
applications. Mater. Today Proc. 2019, 18, 2182–2190. [CrossRef]
3. Schleder, G.R.; Padilha, A.C.; Acosta, C.M.; Costa, M.; Fazzio, A. From DFT to machine learning: Recent approaches to materials
science–a review. J. Phys. Mater. 2019, 2, 032001. [CrossRef]
4. Butler, K.T.; Davies, D.W.; Cartwright, H.; Isayev, O.; Walsh, A. Machine learning for molecular and materials science. Nature
2018, 559, 547–555. [CrossRef]
5. Chibani, S.; Coudert, F.-X. Machine learning approaches for the prediction of materials properties. APL Mater. 2020, 8, 080701.
[CrossRef]
6. Rajendra, P.; Girisha, A.; Naidu, T.G. Advancement of machine learning in materials science. Mater. Today Proc. 2022, 62, 5503–5507.
[CrossRef]
7. Ruoff, R.; Tse, D.S.; Malhotra, R.; Lorents, D.C. Solubility of fullerene (C60) in a variety of solvents. J. Phys. Chem. 1993, 97, 3379–3383.
[CrossRef]
Materials 2023, 16, 5977 26 of 30

8. Guo, K.; Yang, Z.; Yu, C.-H.; Buehler, M.J. Artificial intelligence and machine learning in design of mechanical materials.
Mater. Horiz. 2021, 8, 1153–1172. [CrossRef]
9. Cai, J.; Chu, X.; Xu, K.; Li, H.; Wei, J. Machine learning-driven new material discovery. Nanoscale Adv. 2020, 2, 3115–3130.
[CrossRef]
10. Fang, J.; Xie, M.; He, X.; Zhang, J.; Hu, J.; Chen, Y.; Yang, Y.; Jin, Q. Machine learning accelerates the materials discovery.
Mater. Today Commun. 2022, 33, 104900. [CrossRef]
11. Chen, A.; Zhang, X.; Zhou, Z. Machine learning:Accelerating materials development for energy storage and conversion. InfoMat
2020, 2, 553–576. [CrossRef]
12. Liu, Y.; Niu, C.; Wang, Z.; Gan, Y.; Zhu, Y.; Sun, S.; Shen, T. Machine learning in materials genome initiative: A review. J. Mater.
Sci. Technol. 2020, 57, 113–122. [CrossRef]
13. Raccuglia, P.; Elbert, K.C.; Adler, P.D.; Falk, C.; Wenny, M.B.; Mollo, A.; Zeller, M.; Friedler, S.A.; Schrier, J.; Norquist, A.J.
Machine-learning-assisted materials discovery using failed experiments. Nature 2016, 533, 73–76. [CrossRef] [PubMed]
14. Zhou, L.; Yao, A.M.; Wu, Y.; Hu, Z.; Huang, Y.; Hong, Z. Machine Learning Assisted Prediction of Cathode Materials for Zn-Ion
Batteries. Adv. Theory Simul. 2021, 4, 2100196. [CrossRef]
15. Ridzuan, F.; Zainon, W.M.N.W. A review on data cleansing methods for big data. Procedia Comput. Sci. 2019, 161, 731–738.
[CrossRef]
16. Hossen, M.S. Data preprocess. Machine Learning and Big Data: Concepts, Algorithms, Tools and Applications; Scrivener Publishing:
Beverly, MA, USA, 2020; pp. 71–103.
17. Wu, Y.-W.; Tang, Y.-H.; Tringe, S.G.; Simmons, B.A.; Singer, S.W. MaxBin: An automated binning method to recover individual
genomes from metagenomes using an expectation-maximization algorithm. Microbiome 2014, 2, 26. [CrossRef]
18. Fernández-Delgado, M.; Sirsat, M.S.; Cernadas, E.; Alawadi, S.; Barro, S.; Febrero-Bande, M. An extensive experimental survey of
regression methods. Neural Netw. 2019, 111, 11–34. [CrossRef]
19. Liu, G.-H.; Shen, H.-B.; Yu, D.-J. Prediction of protein–protein interaction sites with machine-learning-based data-cleaning and
post-filtering procedures. J. Membr. Biol. 2016, 249, 141–153. [CrossRef]
20. Wei, J.; Chu, X.; Sun, X.Y.; Xu, K.; Deng, H.X.; Chen, J.; Wei, Z.; Lei, M. Machine learning in materials science. InfoMat 2019,
1, 338–358. [CrossRef]
21. Wang, M.; Wang, T.; Cai, P.; Chen, X. Nanomaterials Discovery and Design through Machine Learning. Small Methods 2019,
3, 1900025. [CrossRef]
22. Schmidt, J.; Marques, M.R.G.; Botti, S.; Marques, M.A.L. Recent advances and applications of machine learning in solid-state
materials science. NPJ Comput. Mater. 2019, 36, 83. [CrossRef]
23. Hou, Y.; Wang, Q.; Tan, T. Prediction of carbon dioxide emissions in China using shallow learning with cross validation. Energies
2022, 15, 8642. [CrossRef]
24. Kurani, A.; Doshi, P.; Vakharia, A.; Shah, M. A comprehensive comparative study of artificial neural network (ANN) and support
vector machines (SVM) on stock forecasting. Ann. Data Sci. 2023, 10, 183–208. [CrossRef]
25. Cover, T.M. Rates of convergence for nearest neighbor procedures. In Proceedings of the Hawaii International Conference on
Systems Sciences, Honolulu, HI, USA, 29–30 January 1968.
26. Peterson, L.E. K-nearest neighbor. Scholarpedia 2009, 4, 1883. [CrossRef]
27. Sharma, A.; Madhushri, P.; Kushvaha, V. Dynamic fracture toughness prediction of fiber/epoxy composites using K-nearest
neighbor (KNN) method. In Handbook of Epoxy/Fiber Composites; Springer: Berlin/Heidelberg, Germany, 2022; pp. 1–16.
28. Sun, B.; Du, J.; Gao, T. Study on the improvement of K-nearest-neighbor algorithm. In Proceedings of the 2009 International
Conference on Artificial Intelligence and Computational Intelligence, Shanghai, China, 7–8 November 2009; pp. 390–393.
29. Hunt, E. Concept Learning: An Information Processing Problem; John Wiley & Sons, Inc.: Hoboken, NJ, USA, 1962.
30. Mak, B.; Munakata, T. Rule extraction from expert heuristics: A comparative study of rough sets with neural networks and ID3.
Eur. J. Oper. Res. 2002, 136, 212–229. [CrossRef]
31. Ruggieri, S. Efficient C4. 5 [classification algorithm]. IEEE Trans. Knowl. Data Eng. 2002, 14, 438–444. [CrossRef]
32. Rokach, L.; Maimon, O. Top-down induction of decision trees classifiers-a survey. IEEE Trans. Syst. Man Cybern. Part C (Appl.
Rev.) 2005, 35, 476–487. [CrossRef]
33. Liu, X.; Liu, T.; Feng, P. Long-term performance prediction framework based on XGBoost decision tree for pultruded FRP
composites exposed to water, humidity and alkaline solution. Compos. Struct. 2022, 284, 115184. [CrossRef]
34. Liu, Y.; Wang, Y.; Zhang, J. New machine learning algorithm: Random forest. In Proceedings of the Information Computing and
Applications: Third International Conference, ICICA 2012, Chengde, China, 14–16 September 2012; pp. 246–252.
35. McCulloch, W.S.; Pitts, W. A logical calculus of the ideas immanent in nervous activity. Bull. Math. Biophys. 1943, 5, 115–133.
[CrossRef]
36. Wu, Y.-C.; Feng, J.-W. Development and application of artificial neural network. Wirel. Pers. Commun. 2018, 102, 1645–1656.
[CrossRef]
37. Huang, Y. Advances in artificial neural networks–methodological development and application. Algorithms 2009, 2, 973–1007.
[CrossRef]
38. Abiodun, O.I.; Jantan, A.; Omolara, A.E.; Dada, K.V.; Mohamed, N.A.; Arshad, H. State-of-the-art in artificial neural network
applications: A survey. Heliyon 2018, 4, e00938. [CrossRef]
Materials 2023, 16, 5977 27 of 30

39. Hmede, R.; Chapelle, F.; Lapusta, Y. Review of neural network modeling of shape memory alloys. Sensors 2022, 22, 5610.
[CrossRef]
40. Mendizabal, A.; Márquez-Neila, P.; Cotin, S. Simulation of hyperelastic materials in real-time using deep learning. Med. Image
Anal. 2020, 59, 101569. [CrossRef] [PubMed]
41. Savaedi, Z.; Motallebi, R.; Mirzadeh, H. A review of hot deformation behavior and constitutive models to predict flow stress of
high-entropy alloys. J. Alloys Compd. 2022, 903, 163964. [CrossRef]
42. Sharma, A.; Madhushri, P.; Kushvaha, V.; Kumar, A. Prediction of the fracture toughness of silicafilled epoxy composites
using K-nearest neighbor (KNN) method. In Proceedings of the 2020 International Conference on Computational Performance
Evaluation (ComPE), Shillong, India, 2–4 July 2020; pp. 194–198.
43. Kumar, P.; Jain, N.K. Surface roughness prediction in micro-plasma transferred arc metal additive manufacturing process using
K-nearest neighbors algorithm. Int. J. Adv. Manuf. Technol. 2022, 119, 2985–2997. [CrossRef]
44. Ghouchan Nezhad Noor Nia, R.; Jalali, M.; Houshmand, M. A Graph-Based k-Nearest Neighbor (KNN) Approach for Predicting
Phases in High-Entropy Alloys. Appl. Sci. 2022, 12, 8021. [CrossRef]
45. Wang, R.; Zheng, Z.; Yin, Z.; Wang, Y. Identification Method of Transformer Winding Material Based on Support Vector Machine.
In Proceedings of the 2022 2nd International Conference on Electrical Engineering and Control Science (IC2ECS), Nanjing, China,
16–18 December 2022; pp. 913–917.
46. Martinez, R.F.; Jimbert, P.; Callejo, L.M.; Barbero, J.I. Material Fracture Life Prediction Under High Temperature Creep Conditions
Using Support Vector Machines And Artificial Neural Networks Techniques. In Proceedings of the 2021 31st International
Conference on Computer Theory and Applications (ICCTA), Alexandria, Egypt, 11–13 December 2021; pp. 127–132.
47. Ahmad, M.; Hu, J.-L.; Ahmad, F.; Tang, X.-W.; Amjad, M.; Iqbal, M.J.; Asim, M.; Farooq, A. Supervised learning methods for
modeling concrete compressive strength prediction at high temperature. Materials 2021, 14, 1983. [CrossRef]
48. Sun, Y.; Li, G.; Zhang, N.; Chang, Q.; Xu, J.; Zhang, J. Development of ensemble learning models to evaluate the strength of
coal-grout materials. Int. J. Min. Sci. Technol. 2021, 31, 153–162. [CrossRef]
49. Samadi, S.H.; Ghobadian, B.; Nosrati, M. Prediction of higher heating value of biomass materials based on proximate analysis
using gradient boosted regression trees method. Energy Sources Part A Recovery Util. Environ. Eff. 2021, 43, 672–681. [CrossRef]
50. Shahmansouri, A.A.; Yazdani, M.; Ghanbari, S.; Bengar, H.A.; Jafari, A.; Ghatte, H.F. Artificial neural network model to predict
the compressive strength of eco-friendly geopolymer concrete incorporating silica fume and natural zeolite. J. Clean. Prod. 2021,
279, 123697. [CrossRef]
51. Liu, Q.-F.; Iqbal, M.F.; Yang, J.; Lu, X.-Y.; Zhang, P.; Rauf, M. Prediction of chloride diffusivity in concrete using artificial neural
network: Modelling and performance evaluation. Constr. Build. Mater. 2021, 268, 121082. [CrossRef]
52. Hinton, G.E.; Salakhutdinov, R.R. Reducing the dimensionality of data with neural networks. Science 2006, 313, 504–507.
[CrossRef] [PubMed]
53. Du, X.; Cai, Y.; Wang, S.; Zhang, L. Overview of deep learning. In Proceedings of the 2016 31st Youth Academic Annual
Conference of Chinese Association of Automation (YAC), Wuhan, China, 11–13 November 2016; pp. 159–164.
54. Agrawal, A.; Choudhary, A. Deep materials informatics: Applications of deep learning in materials science. MRS Commun. 2019,
9, 779–792. [CrossRef]
55. Gu, F.; Khoshelham, K.; Yu, C.; Shang, J. Accurate step length estimation for pedestrian dead reckoning localization using stacked
autoencoders. IEEE Trans. Instrum. Meas. 2018, 68, 2705–2713. [CrossRef]
56. Yang, H.; Shen, S.; Yao, X.; Sheng, M.; Wang, C. Competitive deep-belief networks for underwater acoustic target recognition.
Sensors 2018, 18, 952. [CrossRef]
57. Duong, C.N.; Luu, K.; Quach, K.G.; Bui, T.D. Deep appearance models: A deep boltzmann machine approach for face modeling.
Int. J. Comput. Vis. 2019, 127, 437–455. [CrossRef]
58. Parashar, A.; Raina, P.; Shao, Y.S.; Chen, Y.-H.; Ying, V.A.; Mukkara, A.; Venkatesan, R.; Khailany, B.; Keckler, S.W.; Emer, J.
Timeloop: A systematic approach to dnn accelerator evaluation. In Proceedings of the 2019 IEEE International Symposium on
Performance Analysis of Systems and Software (ISPASS), Madison, WI, USA, 24–26 March 2019; pp. 304–315.
59. Gu, J.; Wang, Z.; Kuen, J.; Ma, L.; Shahroudy, A.; Shuai, B.; Liu, T.; Wang, X.; Wang, G.; Cai, J. Recent advances in convolutional
neural networks. Pattern Recognit. 2018, 77, 354–377. [CrossRef]
60. LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [CrossRef]
61. Wu, S.-w.; Yang, J.; Cao, G.-m. Prediction of the Charpy V-notch impact energy of low carbon steel using a shallow neural network
and deep learning. Int. J. Miner. Metall. Mater. 2021, 28, 1309–1320. [CrossRef]
62. Sun, W.; Li, M.; Li, Y.; Wu, Z.; Sun, Y.; Lu, S.; Xiao, Z.; Zhao, B.; Sun, K. The use of deep learning to fast evaluate organic
photovoltaic materials. Adv. Theory Simul. 2019, 2, 1800116. [CrossRef]
63. Konno, T.; Kurokawa, H.; Nabeshima, F.; Sakishita, Y.; Ogawa, R.; Hosako, I.; Maeda, A. Deep learning model for finding new
superconductors. Phys. Rev. B 2021, 103, 014509. [CrossRef]
64. Li, C.; Wang, C.; Sun, M.; Zeng, Y.; Yuan, Y.; Gou, Q.; Wang, G.; Guo, Y.; Pu, X. Correlated RNN Framework to Quickly Generate
Molecules with Desired Properties for Energetic Materials in the Low Data Regime. J. Chem. Inf. Model. 2022, 62, 4873–4887.
[CrossRef] [PubMed]
65. Zhang, H.; Wang, Y.; Zhao, H.; Lu, K.; Yu, D.; Wen, J. Accelerated topological design of metaporous materials of broadband sound
absorption performance by generative adversarial networks. Mater. Des. 2021, 207, 109855. [CrossRef]
Materials 2023, 16, 5977 28 of 30

66. Unni, R.; Yao, K.; Zheng, Y. Deep convolutional mixture density network for inverse design of layered photonic structures.
ACS Photonics 2020, 7, 2703–2712. [CrossRef]
67. Huang, K.; Lei, H.; Jiao, Z.; Zhong, Z. Recycling waste classification using vision transformer on portable device. Sustainability
2021, 13, 11572. [CrossRef]
68. Jiang, J.; Fan, J.A. Multiobjective and categorical global optimization of photonic structures based on ResNet generative neural
networks. Nanophotonics 2020, 10, 361–369. [CrossRef]
69. Wang, M.; Yang, L.; Zhao, Z.; Guo, Y. Intelligent prediction of wear location and mechanism using image identification based on
improved Faster R-CNN model. Tribol. Int. 2022, 169, 107466. [CrossRef]
70. Ramprasad, R.; Batra, R.; Pilania, G.; Mannodi-Kanakkithodi, A.; Kim, C. Machine learning in materials informatics: Recent
applications and prospects. NPJ Comput. Mater. 2017, 3, 54. [CrossRef]
71. Li, M.; Zhang, H.; Li, S.; Zhu, W.; Ke, Y. Machine learning and materials informatics approaches for predicting transverse
mechanical properties of unidirectional CFRP composites with microvoids. Mater. Des. 2022, 224, 111340. [CrossRef]
72. Ramakrishna, S.; Zhang, T.-Y.; Lu, W.-C.; Qian, Q.; Low, J.S.C.; Yune, J.H.R.; Tan, D.Z.L.; Bressan, S.; Sanvito, S.; Kalidindi, S.R.
Materials informatics. J. Intell. Manuf. 2019, 30, 2307–2326. [CrossRef]
73. Al-Saban, O.; Abdellatif, S.O. Optoelectronic materials informatics: Utilizing random-forest machine learning in optimizing
the harvesting capabilities of mesostructured-based solar cells. In Proceedings of the 2021 International Telecommunications
Conference (ITC-Egypt), Alexandria, Egypt, 13–15 July 2021; pp. 1–4.
74. Zhao, X.-G.; Zhou, K.; Xing, B.; Zhao, R.; Luo, S.; Li, T.; Sun, Y.; Na, G.; Xie, J.; Yang, X. JAMIP: An artificial-intelligence aided
data-driven infrastructure for computational materials informatics. Sci. Bull. 2021, 66, 1973–1985. [CrossRef] [PubMed]
75. Hu, J.; Stefanov, S.; Song, Y.; Omee, S.S.; Louis, S.-Y.; Siriwardane, E.M.; Zhao, Y.; Wei, L. MaterialsAtlas. org: A materials
informatics web app platform for materials discovery and survey of state-of-the-art. NPJ Comput. Mater. 2022, 8, 65. [CrossRef]
76. Kurotani, A.; Kakiuchi, T.; Kikuchi, J. Solubility prediction from molecular properties and analytical data using an in-phase deep
neural network (Ip-DNN). ACS Omega 2021, 6, 14278–14287. [CrossRef]
77. Liang, Z.; Li, Z.; Zhou, S.; Sun, Y.; Yuan, J.; Zhang, C. Machine-learning exploration of polymer compatibility. Cell Rep. Phys. Sci.
2022, 3, 100931. [CrossRef]
78. Zeng, S.; Zhao, Y.; Li, G.; Wang, R.; Wang, X.; Ni, J. Atom table convolutional neural networks for an accurate prediction of
compounds properties. NPJ Comput. Mater. 2019, 5, 84. [CrossRef]
79. Venkatraman, V. The utility of composition-based machine learning models for band gap prediction. Comput. Mater. Sci. 2021,
197, 110637. [CrossRef]
80. Xu, P.; Lu, T.; Ju, L.; Tian, L.; Li, M.; Lu, W. Machine Learning Aided Design of Polymer with Targeted Band Gap Based on DFT
Computation. J. Phys. Chem. B 2021, 125, 601–611. [CrossRef]
81. Espinosa, R.; Ponce, H.; Ortiz-Medina, J. A 3D orthogonal vision-based band-gap prediction using deep learning: A proof of
concept. Comput. Mater. Sci. 2022, 202, 110967. [CrossRef]
82. Wang, T.; Zhang, K.; Thé, J.; Yu, H. Accurate prediction of band gap of materials using stacking machine learning model.
Comput. Mater. Sci. 2022, 201, 110899. [CrossRef]
83. Na, G.S.; Jang, S.; Lee, Y.-L.; Chang, H. Tuplewise material representation based machine learning for accurate band gap prediction.
J. Phys. Chem. A 2020, 124, 10616–10623. [CrossRef] [PubMed]
84. Shen, Z.H.; Liu, H.X.; Shen, Y.; Hu, J.M.; Chen, L.Q.; Nan, C.W. Machine learning in energy storage materials. Interdiscip. Mater.
2022, 1, 175–195. [CrossRef]
85. Feng, Y.; Tang, W.; Zhang, Y.; Zhang, T.; Shang, Y.; Chi, Q.; Chen, Q.; Lei, Q. Machine learning and microstructure design of
polymer nanocomposites for energy storage application. High Volt. 2022, 7, 242–250. [CrossRef]
86. Yue, D.; Feng, Y.; Liu, X.X.; Yin, J.H.; Zhang, W.C.; Guo, H.; Su, B.; Lei, Q.Q. Prediction of Energy Storage Performance in Polymer
Composites Using High-Throughput Stochastic Breakdown Simulation and Machine Learning. Adv. Sci. 2022, 9, 2105773.
[CrossRef] [PubMed]
87. Ojih, J.; Onyekpe, U.; Rodriguez, A.; Hu, J.; Peng, C.; Hu, M. Machine Learning Accelerated Discovery of Promising Thermal
Energy Storage Materials with High Heat Capacity. ACS Appl. Mater. Interfaces 2022, 14, 43277–43289. [CrossRef] [PubMed]
88. Malekloo, A.; Ozer, E.; AlHamaydeh, M.; Girolami, M. Machine learning and structural health monitoring overview with
emerging technology and high-dimensional data source highlights. Struct. Health Monit. 2022, 21, 1906–1955. [CrossRef]
89. Cao, Y.; Miraba, S.; Rafiei, S.; Ghabussi, A.; Bokaei, F.; Baharom, S.; Haramipour, P.; Assilzadeh, H. Economic application of
structural health monitoring and internet of things in efficiency of building information modeling. Smart Struct. Syst. 2020,
26, 559–573.
90. Dang, H.V.; Tatipamula, M.; Nguyen, H.X. Cloud-based digital twinning for structural health monitoring using deep learning.
IEEE Trans. Ind. Inform. 2021, 18, 3820–3830. [CrossRef]
91. Dong, W.; Huang, Y.; Lehane, B.; Ma, G. XGBoost algorithm-based prediction of concrete electrical resistivity for structural health
monitoring. Autom. Constr. 2020, 114, 103155. [CrossRef]
92. Fu, B.; Chen, S.-Z.; Liu, X.-R.; Feng, D.-C. A probabilistic bond strength model for corroded reinforced concrete based on weighted
averaging of non-fine-tuned machine learning models. Constr. Build. Mater. 2022, 318, 125767. [CrossRef]
93. Gao, J.; Koopialipoor, M.; Armaghani, D.J.; Ghabussi, A.; Baharom, S.; Morasaei, A.; Shariati, A.; Khorami, M.; Zhou, J. Evaluating
the bond strength of FRP in concrete samples using machine learning methods. Smart Struct. Syst. Int. J. 2020, 26, 403–418.
Materials 2023, 16, 5977 29 of 30

94. Li, Z.; Qi, J.; Hu, Y.; Wang, J. Estimation of bond strength between UHPC and reinforcing bars using machine learning approaches.
Eng. Struct. 2022, 262, 114311. [CrossRef]
95. Su, M.; Zhong, Q.; Peng, H.; Li, S. Selected machine learning approaches for predicting the interfacial bond strength between
FRPs and concrete. Constr. Build. Mater. 2021, 270, 121456. [CrossRef]
96. Khan, B.M.; Cohen, Y. Predictive Nanotoxicology: Nanoinformatics Approach to Toxicity Analysis of Nanomaterials. In Machine
Learning in Chemical Safety and Health: Fundamentals with Applications; John Wiley & Sons: Hoboken, NJ, USA, 2022; pp. 199–250.
97. Huang, Y.; Li, X.; Cao, J.; Wei, X.; Li, Y.; Wang, Z.; Cai, X.; Li, R.; Chen, J. Use of dissociation degree in lysosomes to predict metal
oxide nanoparticle toxicity in immune cells: Machine learning boosts nano-safety assessment. Environ. Int. 2022, 164, 107258.
[CrossRef] [PubMed]
98. Gousiadou, C.; Marchese Robinson, R.; Kotzabasaki, M.; Doganis, P.; Wilkins, T.; Jia, X.; Sarimveis, H.; Harper, S. Machine learning
predictions of concentration-specific aggregate hazard scores of inorganic nanomaterials in embryonic zebrafish. Nanotoxicology
2021, 15, 446–476. [CrossRef] [PubMed]
99. Liu, L.; Zhang, Z.; Cao, L.; Xiong, Z.; Tang, Y.; Pan, Y. Cytotoxicity of phytosynthesized silver nanoparticles: A meta-analysis by
machine learning algorithms. Sustain. Chem. Pharm. 2021, 21, 100425. [CrossRef]
100. Sajid, M.; Ihsanullah, I.; Khan, M.T.; Baig, N. Nanomaterials-based adsorbents for remediation of microplastics and nanoplastics
in aqueous media: A review. Sep. Purif. Technol. 2022, 305, 122453. [CrossRef]
101. Moosavi, S.; Manta, O.; El-Badry, Y.A.; Hussein, E.E.; El-Bahy, Z.M.; Mohd Fawzi, N.f.B.; Urbonavičius, J.; Moosavi, S.M.H.
A study on machine learning methods’ application for dye adsorption prediction onto agricultural waste activated carbon.
Nanomaterials 2021, 11, 2734. [CrossRef]
102. Guo, W.; Liu, J.; Dong, F.; Chen, R.; Das, J.; Ge, W.; Xu, X.; Hong, H. Deep learning models for predicting gas adsorption capacity
of nanomaterials. Nanomaterials 2022, 12, 3376. [CrossRef]
103. Gómez-Bombarelli, R.; Aguilera-Iparraguirre, J.; Hirzel, T.D.; Duvenaud, D.; Maclaurin, D.; Blood-Forsythe, M.A.; Chae, H.S.;
Einzinger, M.; Ha, D.-G.; Wu, T. Design of efficient molecular organic light-emitting diodes by a high-throughput virtual screening
and experimental approach. Nat. Mater. 2016, 15, 1120–1127. [CrossRef]
104. Xue, D.; Yuan, R.; Zhou, Y.; Xue, D.; Lookman, T.; Zhang, G.; Ding, X.; Sun, J. Design of high temperature Ti-Pd-Cr shape memory
alloys with small thermal hysteresis. Sci. Rep. 2016, 6, 28244. [CrossRef] [PubMed]
105. Li, W.; Yang, T.; Liu, C.; Huang, Y.; Chen, C.; Pan, H.; Xie, G.; Tai, H.; Jiang, Y.; Wu, Y. Optimizing piezoelectric nanocomposites by
high-throughput phase-field simulation and machine learning. Adv. Sci. 2022, 9, 2105550. [CrossRef] [PubMed]
106. Zhang, L.; He, M.; Shao, S. Machine learning for halide perovskite materials. Nano Energy 2020, 78, 105380. [CrossRef]
107. Li, L.; Tao, Q.; Xu, P.; Yang, X.; Lu, W.; Li, M. Studies on the regularity of perovskite formation via machine learning. Comput. Mater.
Sci. 2021, 199, 110712. [CrossRef]
108. Liu, H.; Cheng, J.; Dong, H.; Feng, J.; Pang, B.; Tian, Z.; Ma, S.; Xia, F.; Zhang, C.; Dong, L. Screening stable and metastable ABO3
perovskites using machine learning and the materials project. Comput. Mater. Sci. 2020, 177, 109614. [CrossRef]
109. Omprakash, P.; Manikandan, B.; Sandeep, A.; Shrivastava, R.; Viswesh, P.; Panemangalore, D.B. Graph representational learning
for bandgap prediction in varied perovskite crystals. Comput. Mater. Sci. 2021, 196, 110530. [CrossRef]
110. Wang, Z.; Cai, J.; Wang, Q.; Wu, S.; Li, J. Unsupervised discovery of thin-film photovoltaic materials from unlabeled data.
NPJ Comput. Mater. 2021, 7, 128. [CrossRef]
111. Huang, K.; Zhan, X.-L.; Chen, F.-Q.; Lü, D.-W. Catalyst design for methane oxidative coupling by using artificial neural network
and hybrid genetic algorithm. Chem. Eng. Sci. 2003, 58, 81–87. [CrossRef]
112. Zhang, S.; Lu, S.; Zhang, P.; Tian, J.; Shi, L.; Ling, C.; Zhou, Q.; Wang, J. Accelerated Discovery of Single-Atom Catalysts for
Nitrogen Fixation via Machine Learning. Energy Environ. Mater. 2023, 6, e12304. [CrossRef]
113. Wei, S.; Baek, S.; Yue, H.; Liu, M.; Yun, S.J.; Park, S.; Lee, Y.H.; Zhao, J.; Li, H.; Reyes, K. Machine-learning assisted exploration:
Toward the next-generation catalyst for hydrogen evolution reaction. J. Electrochem. Soc. 2021, 168, 126523. [CrossRef]
114. Hueffel, J.A.; Sperger, T.; Funes-Ardoiz, I.; Ward, J.S.; Rissanen, K.; Schoenebeck, F. Accelerated dinuclear palladium catalyst
identification through unsupervised machine learning. Science 2021, 374, 1134–1140. [CrossRef] [PubMed]
115. Zhang, J.; Zhu, Z.; Xiang, X.-D.; Zhang, K.; Huang, S.; Zhong, C.; Qiu, H.-J.; Hu, K.; Lin, X. Machine learning prediction of
superconducting critical temperature through the structural descriptor. J. Phys. Chem. C 2022, 126, 8922–8927. [CrossRef]
116. Le, T.D.; Noumeir, R.; Quach, H.L.; Kim, J.H.; Kim, J.H.; Kim, H.M. Critical temperature prediction for a superconductor: A
variational bayesian neural network approach. IEEE Trans. Appl. Supercond. 2020, 30, 8600105. [CrossRef]
117. Zhang, J.; Zhang, K.; Xu, S.; Li, Y.; Zhong, C.; Zhao, M.; Qiu, H.-J.; Qin, M.; Xiang, X.-D.; Hu, K. An integrated machine learning
model for accurate and robust prediction of superconducting critical temperature. J. Energy Chem. 2023, 78, 232–239. [CrossRef]
118. Roter, B.; Dordevic, S. Predicting new superconductors and their critical temperatures using machine learning. Phys. C Supercond.
Its Appl. 2020, 575, 1353689. [CrossRef]
119. Pereti, C.; Bernot, K.; Guizouarn, T.; Laufek, F.; Vymazalová, A.; Bindi, L.; Sessoli, R.; Fanelli, D. From individual elements to
macroscopic materials: In search of new superconductors via machine learning. NPJ Comput. Mater. 2023, 9, 71. [CrossRef]
120. Xie, Y.; Zhang, C.; Hu, X.; Zhang, C.; Kelley, S.P.; Atwood, J.L.; Lin, J. Machine learning assisted synthesis of metal–organic
nanocapsules. J. Am. Chem. Soc. 2019, 142, 1475–1481. [CrossRef]
Materials 2023, 16, 5977 30 of 30

121. Pellegrino, F.; Isopescu, R.; Pellutiè, L.; Sordello, F.; Rossi, A.M.; Ortel, E.; Martra, G.; Hodoroaba, V.-D.; Maurino, V. Machine
learning approach for elucidating and predicting the role of synthesis parameters on the shape and size of TiO2 nanoparticles.
Sci. Rep. 2020, 10, 18910. [CrossRef]
122. Tao, H.; Wu, T.; Aldeghi, M.; Wu, T.C.; Aspuru-Guzik, A.; Kumacheva, E. Nanoparticle synthesis assisted by machine learning.
Nat. Rev. Mater. 2021, 6, 701–716. [CrossRef]
123. Braham, E.J.; Cho, J.; Forlano, K.M.; Watson, D.F.; Arròyave, R.; Banerjee, S. Machine learning-directed navigation of synthetic
design space: A statistical learning approach to controlling the synthesis of perovskite halide nanoplatelets in the quantum-
confined regime. Chem. Mater. 2019, 31, 3281–3292. [CrossRef]
124. Epps, R.W.; Bowen, M.S.; Volk, A.A.; Abdel-Latif, K.; Han, S.; Reyes, K.G.; Amassian, A.; Abolhasani, M. Artificial chemist: An
autonomous quantum dot synthesis bot. Adv. Mater. 2020, 32, 2001626. [CrossRef] [PubMed]
125. Wang, J.; Wang, Y.; Chen, Y. Inverse design of materials by machine learning. Materials 2022, 15, 1811. [CrossRef] [PubMed]
126. Li, S.; Barnard, A.S. Inverse Design of Nanoparticles Using Multi-Target Machine Learning. Adv. Theory Simul. 2022, 5, 2100414.
[CrossRef]
127. Wang, R.; Liu, C.; Wei, Y.; Wu, P.; Su, Y.; Zhang, Z. Inverse design of metal nanoparticles based on deep learning. Results Opt.
2021, 5, 100134. [CrossRef]
128. He, J.; He, C.; Zheng, C.; Wang, Q.; Ye, J. Plasmonic nanoparticle simulations and inverse design using machine learning.
Nanoscale 2019, 11, 17444–17459. [CrossRef]

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual
author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to
people or property resulting from any ideas, methods, instructions or products referred to in the content.

Common questions

Powered by AI

ML can reduce computational costs, shorten development cycles, and improve computational accuracy, making it a promising approach for novel materials screening and predicting material properties. It provides guidance for stable and efficient material synthesis while offering novel ideas and research methodologies for materials science .

The ip-DNN model brought innovation to solubility prediction by using a multi-step approach, predicting intermediate molecular components and descriptors first, to enhance accuracy. It analyzed input data like NMR information, refractive index, and density, and used intermediate regression models to refine predictions, demonstrating a unique methodology in solubility prediction .

The major challenges include the requirement for large amounts of data, which are difficult to obtain due to high acquisition costs and lack of standardized data processing. Moreover, the datasets often suffer from excessive concentration or dispersion. The selection of appropriate features is critical but challenging, and manual feature engineering can overlook significant details, whereas automated feature engineering offers a possible solution. Additionally, ML models are often 'black boxes,' necessitating experimental verification and further study of the underlying physicochemical laws .

Liang et al. developed a generalized ML method using artificial neural networks (ANNs) to predict polymer compatibility, which is the total miscibility of polymers at the molecular scale. They built a database by aggregating data from scattered literature using natural language processing techniques. The method allowed predictions based on the molecular structure of the polymers and their compositions, achieving a prediction accuracy of at least 75% on a dataset with 1400 entries .

The use of "black-box" ML models implies a lack of interpretability, which necessitates experimental verification of results and a deep understanding of underlying physicochemical laws. While they offer potential in exploratory tasks and can guide synthesis and property prediction, traditional methods remain critical due to ML models' current limitations in explanation and understanding .

Machine learning accelerates the discovery of single-atom catalysts for nitrogen fixation by quickly analyzing vast datasets to identify promising catalyst candidates. This method leverages ML's ability to process and evaluate extensive materials data efficiently, bypassing the slow, traditional trial-and-error experimental approaches .

Automated feature engineering is beneficial because it can automatically construct new candidate features and select those most appropriate for model training, effectively addressing challenges where manual feature selection might be biased by the researcher's intuition or experience, potentially overlooking important features .

The atomic table CNN model predicts material properties such as the band gap and ground energy with greater accuracy compared to traditional DFT calculations. This model not only provides accurate results but does so with lower computational costs, highlighting ML's potential to outperform traditional methods in specific applications within material science .

The potential future of ML in materials science is tremendous, with prospects of eventually replacing traditional research methods by bridging exploratory tasks and offering innovative research strategies. Further theoretical and methodological improvements might enable ML to take on more comprehensive roles in materials prediction, discovery, and synthesis, beyond its current exploratory capacity .

Data preprocessing and feature engineering enhance the dataset's structure, enabling computers to better understand physicochemical relationships of materials, thereby improving the detection and prediction of material properties. High-quality data is crucial for the effective functioning of ML models, as the final results highly depend on the data's amount and reliability .

You might also like