0% found this document useful (0 votes)
8 views21 pages

Example Standard References

The document outlines the importance of sustainable ceramics and the role of data science in enhancing their production processes. It discusses the energy-intensive nature of ceramics manufacturing, which contributes significantly to CO2 emissions, and proposes a data-centered approach to improve sustainability through machine learning and comprehensive databases. The literature review highlights successful applications of data science in ceramics research, emphasizing the need for reliable property databases to facilitate the transition to more sustainable practices.

Uploaded by

student631
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views21 pages

Example Standard References

The document outlines the importance of sustainable ceramics and the role of data science in enhancing their production processes. It discusses the energy-intensive nature of ceramics manufacturing, which contributes significantly to CO2 emissions, and proposes a data-centered approach to improve sustainability through machine learning and comprehensive databases. The literature review highlights successful applications of data science in ceramics research, emphasizing the need for reliable property databases to facilitate the transition to more sustainable practices.

Uploaded by

student631
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

1.

Standard citation format is as follows:


References are to be numbered, ordered sequentially as they appear in the
text. References should appear as superscript numbers, before the
punctuation mark, and without a space between the superscript number and
the word. Only one publication should be listed for each number.
Examples:
Single citation:
These conduction electrons can oscillate and give rise to surface plasmon
resonance1. Surface plasmon resonance is a technique used to characterize
binding interactions2.
Multiple citations:
These conduction electrons can oscillate and give rise to surface plasmon
resonance1,2. Surface plasmon resonance is a technique used to characterize
binding interactions3.

Ceramics Research Done Smartly: Accelerating the Transition to Sustainable


Ceramics

0. Abstract
Ceramics have been a witness of human ingenuity throughout history. From their usage in
pottery and artifacts in the past to industrial, electrical, and aerospace applications of ceramics in
the present, ceramics are ubiquitous in our lives. Unfortunately, manufacturing ceramics is
extremely energy-intensive and is responsible for more than 8% of global anthropogenic CO2
emissions. There have been continuous efforts to make these valuable materials more
sustainable. Recently, with the advance of data science, organizing and learning from data has
become much less time-consuming. This literature review aims to coalesce the two fields,
sustainability, and data science, with a focus on ceramics. Instead of narrowing our focus to
specific ceramics, such as glass or cement, we proposed procedures that can be broadly applied.
By analyzing different data science approaches and their applicability to creating ceramics
databases, we evaluated various ways to promote ceramics sustainability using data science.
More specifically, we analyzed vital insights or techniques that may be applied elsewhere and
sought opportunities to create procedures that required the least human input and benefited the
most from machine learning models. We categorized different approaches toward regression and
data extraction, examined suitable approaches for various scenarios, and qualified these efforts
by explaining the impact new databases can have on the sustainability of ceramics. We hope this
literature review will prove relevant to readers with an interest in data science, sustainable
ceramics, and their intersections.

Keywords: Ceramics, Sustainability, Data Science, Machine Learning (ML), Deep Learning,
Large Language Model (LLM), Regression, Data Extraction

1. Introduction
Ceramics are defined as materials that are neither metallic nor organic. These materials often
present themselves with valuable properties, such as heat resistance, semiconducting behavior, or
inertness. For example, refractories are used in extreme-temperature industrial processes because
of their high resistance to heat, and technical ceramics are used in chemical processing plants
because of their chemical inertness1. Ceramics are undoubtedly valuable, but unfortunately,
producing ceramics is a highly energy-intensive process. Aside from the mechanical energy
needed to mill and pulverize the mined raw materials (e.g., rocks of clays, silicas, limestones),
these raw materials require incredibly high temperatures (e.g., above 1500 °C in specially
equipped furnaces) to react in solid or molten form and yield the solid solutions that define the
products. The associated greenhouse gas emissions are not attributable only to the energy input
to the manufacturing process (e.g., process heat, electricity, mechanical power) but also to the
process emissions emanating from the solid-state reactions responsible for chemically converting
the raw materials to products, accentuating the climate impact of the ceramics industry via these
process emissions. For instance, emissions from the European ceramic industry amount to 19
million tons of CO2 annually1, and much research is necessary to make the ceramics industry
more sustainable.

We provide a general schematic of the ceramics industry below to better explain the
sustainability challenge in ceramics. After mining from Earth deposits, raw materials such as
clay, silica, or limestone are transported to and stored in the manufacturing plant. This
transportation process could lead to significant emissions if raw materials are mined far from the
ceramic plant. The raw materials at the ceramic plant are then milled, mixed, and pulverized to
increase the surface area available for the ensuing solid-state reactions. Often, they get formed
into a paste and shaped into a mold before getting dried, to prepare them for the firing step. The
raw materials then react under firing in furnaces, where heat management and distribution are of
paramount importance to the quality of the product. Afterward, finishing often includes cooling,
packaging, and regular testing for quality control of the products.
Figure 1: A Simplified General Schematic of the Ceramics Manufacturing Process adapted from Ceramic Roadmap
to 2050
Figure 1 overlays some examples of opportunities for sustainability in existing ceramic plants.
Following the schematic above, one can start by placing the ceramics manufacturing site near the
mining and raw materials extraction site to minimize transportation emissions. The drying step is
crucial before firing, as this step controllably gets rid of water and volatile organic impurities
contaminating the raw materials, ensuring the mechanical integrity of the fragile and, as of yet,
wet shapes. Here, there exists an opportunity for heat integration, where the flue gases from the
firing step heat the drying air or are applied directly to dry the incoming feed to the furnace.
Using environmentally sustainable fuel sources such as green hydrogen during this firing stage
could significantly improve the sustainability of the overall process. The furnace is also lined
with ceramic refractories that trap the heat inside, offering effective insulation. Further, capturing
the carbon released during the drying and firing stages can help decarbonize the production
process. However, further research must be done to find more practical carbon removal
technologies and offsetting measures because current methods, like carbon capture and storage,
are too costly to be practical. The ceramics industry is heading towards a circular pathway as
well, in which refuse from manufacturing and distribution sites is used as a filler to minimize
new raw materials in the production mix, solving both a waste problem and an emissions
problem. Some wall and floor tiles industry manufacturers transitioned from rotary printing to
digital printing and created walls and floor tiles with 80% recycled ceramics using ceramic inks
instead of decorative pastes. Clay pipes consisting of 40% new raw materials and refractories
containing 20% to 80% new raw materials can be produced, and it is possible to reuse 90% of
expanded clay products1. The goal should be to make all ceramic products 100% reusable so no
waste is created. Although this might require significant research or technological breakthroughs,
decreasing the percentage of raw materials used would result in substantial reductions in process
emissions.

However, there is only so much that such improvements can do to turn the ceramics industry into
a genuine engine for sustainability, for the ceramics industry remains intrinsically emissions-
intensive and geared towards end products rather than advanced ceramics that go into further
green technologies1. Here, we suggest that a smart approach towards sustainable ceramics
research is data-centered, including building, maintaining, and seamlessly interacting with
comprehensive databases that adhere to FAIR data principles (findable, accessible, interoperable,
and reusable).

To make this point, we highlight two recent articles discussing the highly successful results of
this endeavor. The first is by K. Gong and coworkers at Princeton University, where they
established a strong research program geared towards the manufacture of cement with less
emissions, by designing alkali-activated kaolin clays as raw materials that help turn industrial
wastes and already-calcined clays (which emit no further CO2 upon processing) into cementitious
binders, minimizing process emissions2. Specifically, they applied computationally intensive
density functional theory simulations to calculate the binding energies of seventeen such alkali
cations to aluminosilicate dimers and trimers and then correlated these binding energies to two
other properties that are readily obtainable (i.e., do not require DFT simulations): the ionic
potential and field strength of these cations. The high-quality second-order polynomial
correlations (R2=0.99-1.00) highlight their usability to rapidly estimate such binding energies of
new cations or new substrates, extending this research's impact into materials not initially
simulated by the authors2. The data on binding energies of different cations to different clay
materials can help design alkali-activated kaolin clays as raw materials by providing insight into
the performances (e.g., performance against sulfuric acid attacks) of the resulting sustainable
cement. This endeavor highlights the importance of a data-centered approach toward innovative
ceramics research.

The second article is by S. Kim and coworkers at Seoul National University. They successfully
created an extensive band-gap database for semiconducting inorganic materials, as applied in
photovoltaic devices. By using a hybrid functional and considering a stable magnetic ordering,
they made a database for band gaps for 10,841 materials with a significantly smaller root-mean-
square error of 0.36 eV than the 0.75-1.05 eV of existing databases (1 eV ~ 100 kJ/mol)3.
Additionally, they identified a considerable number of small-gap materials that were
misclassified as metals in other databases. By correctly classifying these materials as
semiconducting inorganics, they provided a broader selection of materials that manufacturers or
researchers may consider using. This can come with various benefits, such as reduced costs of
materials, reduced transport emissions, and more. Overall, S. Kim and coworkers created a
highly impactful (and highly cited) database utilizing a data-centered approach, further
highlighting the potential of a data-centered approach for scientific research and innovation.
Band gaps are but one of a myriad of properties that define the utility of ceramic materials in
advanced and sustainable applications (see Table 1 for a non-comprehensive list of such
properties and applications).
Table 1: A Simplified Subset of Major Property Categories for Technical and Sustainable Ceramics

Property Definition Application Type Unit

Thermal Ability to Insulation, Thermal W/m/K (Watts


Conductivity conduct heat electronic per meter-
cooling Kelvin)

Band Gap Energy Semiconductors, Electronic eV (electron


difference solar cells, LEDs volt)
between the
highest valence
bond and the
lowest
conduction band

Piezoelectricity Ability to Sensors, Electrical/ Vm/N (voltage-


generate an actuators, Mechanical meter per
electric charge in transducers Newton)
response to
mechanical
stress, or vice
versa

Refractive Index Measure of light Lenses, photonic Optical dimensionless


bending in a devices, optical
material networks

Process Emissions Ceramic Manufacturing kg CO2/kg


Emissions generated during (cement) product
manufacturing or industry (kilogram CO2
industrial per kilogram
processes product)

Corrosion Ability to Chemical Chemical mm/yr


Resistance withstand processing (millimeter per
degradation by year)
chemical
reactions

Melting Point Temperature at Material Thermal °C or K (degrees


which a material processing, Celsius or
changes from casting, welding Kelvin)
solid to liquid
Fracture Ability to resist Structural Mechanical MPa√ m
Toughness fracture when a ceramics, dental (megapascal
crack is present implants times the square
root of meters)

Oxygen Ionic Ability to Solid oxide fuel Electrical S/m (siemens per
Conductivity conduct oxygen cells, oxygen meter)
ions under an sensors
applied voltage
gradient

Lamentably, there is a severe lack of comprehensive, reliable, up-to-date, and accessible


ceramics property databases4, a problem that must be addressed to transition to sustainable
ceramics efficiently. This paper suggests combinations of already-present machine learning
methods to maximize the accuracy of ceramics property databases that can be extracted with
minimal human input, current literature, and available data.

2. Results
2.1. An Overview of Machine Learning Models
Following the recent rise of machine learning, there have been recent efforts to extract data from
text using natural language processing models. For example, M. Polak and coworkers proposed a
successful procedure using ChatGPT to extract mid-sized databases from text5. This process
requires minimal prior knowledge of coding and the property being analyzed, but can generate
databases with high precision and recall in a single day. This automatic quality of machine
learning, being able to extract data from literature or learn from available data to create
sufficiently accurate databases with minimal human input, is what makes machine learning so
valuable and, thus, will be the focus of this literature review.

In the world of data science, machine learning methods have varying types of training data and
tasks to solve. Supervised models or algorithms refer to models trained on data sets with labels
— either categorical or continuous — corresponding to the data points, and can help predict
material properties. Some supervised models, such as graph neural networks, can input graphs as
data, which is useful when predicting properties based on molecular structures that are
represented as graphs. In contrast, unsupervised models are trained on data sets that lack such
labels and are often used for clustering or anomaly detections, but lack utility for the purpose of
creating databases for ceramics data. This paper will focus on the following three most common
types of supervised machine learning models that were present in the literature surveyed:
Support Vector Machines, Decision Trees, and Neural Networks. A brief overview of these three
types of models is described in Figure 2.
Figure 2: A Simple Overview of Three Types of Machine Learning Models

2.1.1. Support Vector Machines


Support Vector Machines (SVMs) are supervised learning algorithms that can be applied for
both classification and regression. For classification tasks, the machine finds the most optimal
decision boundary, referred to as a hyperplane, that maximizes the margin, or distance, between
the nearest data points of two different classes and the decision boundary. A larger margin
implies a clearer separation between the two classes of data and, therefore, a better generalized
classifying model. An SVM for regression, often referred to as Support Vector Regression
(SVR), is an extension of SVM that predicts continuous variables. Instead of finding a decision
boundary that maximizes the margin, an SVR finds a function that predicts the output value
within a specified margin of tolerance6.

2.1.2. Decision Trees


Decision Trees are supervised learning algorithms that utilize tree-like structures to make
predictions. These trees are composed of nodes and branches connecting the nodes. The starting
node at the top of the tree is called a root node, where the input is given. The nodes at the bottom
of the tree are called leaf nodes, and these nodes represent all possible outcomes within the
dataset or scenario. Starting from the root node, at each node, the machine makes a decision
based on its training of previous data points and sends the input to a new node that was
connected to the previous node via a branch. This process is then repeated until the input arrives
at a terminal node representing a possible output. There are various types of decision trees,
including Random Forest, Gradient Boost, Extreme Gradient Boost, and Adaptive Boost decision
trees, all of which are commonly used for regression. Figure 3 below is a brief overview of these
four special types of decision trees for regression7 8 9.

Figure 3: A Non-Comprehensive Diagram Depicting Various Decision Trees for Regression

A random forest decision tree is an ensemble of decision trees trained separately on random
samples of the training data. These independent decision trees are run in parallel when given an
input, and the individual results generated by each of these decision trees are aggregated to make
a final prediction. A gradient boost decision tree is a series of several decision trees that attempt
to predict the errors from the previous decision trees in the series. An extreme gradient boost
decision tree is a gradient boost decision tree with regularizations that prevent complex models
that are overfitting the training data and/or aren’t generalizing well. These regularizations
include but are not limited to, lasso regularizations, which encourage sparsity, and ridge
regularizations, which reduce the variance of the model. An adaptive boost decision tree consists
of a group of independent decision trees that learn adaptively, forcing decision trees that made
significant errors to train more on those instances to reduce the large errors. The results from
these independent decision trees are then weighted and compiled so that the decision trees that
made those significant errors have less influence over the result7 8 9.

2.1.3. Neural Networks


A Neural Network is a program or model that intends to replicate how our brain makes
decisions. The nodes in a neural network are connected to one another and serve as neurons of
the machine. The individual nodes input data, calculate a weighted output using the input data,
and compare the output to a specified threshold value to “make decisions”10. When these models
are trained on large data samples, they can classify, cluster, and estimate data rapidly and with
high accuracy. Figure 4 is a diagram depicting common types of neural networks and a non-
comprehensive type of tasks these models can complete.

Figure 4: A Non-Comprehensive Diagram Depicting Types of Neural Networks and Tasks

Graph Neural Networks are a type of neural network that can intake graph-structured irregular
data. These neural networks are especially useful in understanding materials because they can
input molecular structures as data and handle them efficiently. Deep Neural Networks,
commonly referred to as Deep Learning, are neural networks with three or more layers, which is
common for models that require complexity. These models can be trained to estimate material
properties using other more available properties without knowing the specific relation between
those properties beforehand. A Large Language Model (LLM) is an application of deep learning
that utilizes multiple layers of deep neural networks for natural language processing and is
trained on extremely large samples of textual data. These various types of neural networks can be
useful in ceramics research because of their ability to interpret large amounts of data or text, to
predict numerical data, or to extract different data from the vast sea of literature 11.
2.2. Machine Learning Models Used for Regression
Machine learning models such as deep neural networks and the various decision trees can help
solve regression tasks by using data that are more easily accessible to predict values of properties
that are expensive or computationally intensive to compute. This is a common practice for
scientific researchers, as seen by how K. Gong and coworkers successfully calculated binding
energies using density functional theory (DFT) and then extrapolated them with second-order
polynomials to avoid further DFT calculations on metal cations not previously calculated 2. The
main benefit of using machine learning models to make these predictions is that they can
calculate predictions nearly instantaneously with no human input or background information on
these properties.

2.2.1. Graph Neural Networks Used with Molecular Structure Data


The benefits of utilizing neural networks to estimate materials property data become apparent
when considering interatomic potentials or inferring properties of materials that require DFT but
are prohibitively expensive to obtain large samples of (e.g., high-entropy materials or catalysts).
Graph neural networks specifically serve as relevant tools to predict these properties because of
their ability to input the molecular structures of these high-entropy materials and predict
properties rapidly. Such was the case when C. M. Clausen and coworkers successfully performed
zero-shot inferences of adsorption energies of high-entropy materials utilizing a graph neural
network. Here, zero-shot inference refers to the direct prediction of relaxed structure models and
adsorption energies of materials using a machine learning model that was not explicitly trained to
output those relaxed structures given an initial structure. Their model possessed inference speeds
that were “practically instantaneous” thanks to the low number of parameters used in the neural
network. Despite already achieving “state-of-the-art accuracy,” yielding mean absolute errors in
adsorption energies of 0.04 eV (adsorption energies of these high-entropy materials are normally
on the order of magnitude of 100), with the fraction of the time DFT would take to emulate the
high-entropy materials, C. M. Clausen and coworkers stated that a hyperparameter optimization
could “further improve the performance”12. C. M. Clausen and coworkers’ work highlights the
possibilities of using machine learning to accurately estimate data on materials that were out of
the domain of previous experiments or data collections. Optimizing hyperparameters and
increasing the depth of the neural networks appear to be valuable steps to take in the future
because of the amount of time and resources these machine-learning models can save when
compared to running DFT calculations with supercomputers. Further, substituting machine
learning models in place of DFT calculations allows research normally involving computational
quantum mechanical modeling methods to be much more accessible to researchers around the
world, as supercomputers for scientific research may not be readily accessible for researchers,
especially those in less developed countries.
2.2.2. Decision Trees Used with Continuous Data
Although the nature of decision trees makes them seem suitable for classification tasks related to
categorial variables, some types of decision trees can be powerful tools that can predict
quantitative variables rapidly. Similar to the application of graph neural networks, decision trees
can be used to estimate properties of materials that are on par with the quality obtained from
DFT calculations and experimental results. For example, K. Kaufmann and coworkers were able
to use random forest decision trees to estimate the synthesizability of high-entropy materials 13. K.
Kaufmann and coworkers first chose a set of around 800 attributes of materials that may be
relevant for synthesizability to feed as features for the machine learning model. Then, they
utilized a feature named the SelectFromModel in Scikit-learn14, a machine learning tool in
Python, to narrow down the eight most relevant features to be used to train their model. Among
these eight features related to chemistry, physics, or thermodynamics, K. Kaufmann and
coworkers were able to predict which of these features contributed the most towards the final
prediction of synthesizability13. This is a strong benefit of using decision trees for regression
tasks: allowing us to peek inside the black box of machine learning models. Thanks to its
flowchart-like structure, decision tree models are one of the most interpretable machine learning
models available. In scenarios like these, where some sort of connection between quantitative
variables is expected, being able to predict which quantitative attributes were responsible for the
data can bring researchers one step closer to understanding the underlying cause and science that
they were unaware of, which is why the interpretable quality of decision trees is valuable.

Unfortunately, some decision tree algorithms are much more complicated in nature and are
harder to interpret than random forest decision trees or regular decision trees. Such is the case for
extreme gradient boost and adaptive boost decision trees. Because normalizations and
optimizations are applied to extreme gradient boost trees and because adaptive boost decision
trees compare the training data, they are intrinsically too complex to interpret or understand.
Still, these decision trees take comparatively much less time to make predictions. For example,
when J. Deb and coworkers predicted the ablation performance of ceramic matrix composites
using various decision trees to compare results, the extreme gradient boost and adaptive boost
decision trees took significantly less simulation time7. Specifically, when using the performance
parameters that resulted in the greatest R-squared scores and the least mean absolute errors
among all decision trees, the extreme gradient boost and adaptive boost decision trees took 1
second and 2 seconds accordingly, the random forest decision tree took slightly longer than 2
minutes, the regular decision tree took longer than 18 minutes, and the gradient boost decision
tree took close to 9 minutes7. This shows that the decision trees that are too complicated to be
interpreted still serve a purpose, as they can drastically improve computing efficiency while still
retaining high accuracy.
2.2.3. Flowchart of Machine Learning Methods to Utilize for Different
Regression Tasks

Figure 5: A Brief Overview of What Machine Learning Models to Use Given Different Scenarios

The flowchart depicted in Figure 5 shows a summary of what machine learning models one
should use depending on the type of regression task assigned. The properties of ceramics are
dependent on the molecular structure of the material. As such, it is plausible that some form of
molecular structure data is involved with the inputs or outputs of a given regression task. In such
cases, it is encouraged to use a deep neural network to be able to take full advantage of the
structured data. Further, predicting molecular structures of materials as an intermediate step
before predicting a quantitative property could be useful, as knowing the molecular structures
can help determine the distribution of electron density across the material, find distances between
atoms, and find angles between bonds, all of which allow us to utilize DFT or quantum
mechanical modeling to predict other properties of a material, if such level of accuracy is
necessary for a project or experiment. For example, a graph neural network can be used to help
narrow down which candidate has the significance to further test using expensive DFT
calculations to make more accurate results.

When predicting continuous, quantitative data, you can utilize various decision trees or a support
vector regression. In particular, the extreme gradient boost decision tree, adaptive boost decision
tree, and random forest decision tree models execute the fastest among these six machine
learning models shown in Figure 5. This was shown experimentally in J. Deb and coworkers’
experiments7, and aligns with the properties of the model: the extreme gradient boost tree is an
optimization of a gradient boost tree that implements parallelization to speed up the tree-building
process, while the adaptive boost decision tree and the random forest decision trees generate
small trees to aggregate their results, which takes less time than generating a larger, more
complex tree. Other than the apparent difference in the expected execution time of these models,
it is hard to predict which model will perform the best in terms of accuracy or precision because
of how unique these models are. Here, it is worth noting that the comparative performance of
these models remains relatively constant even as the test data increases. Despite the small
changes in the performances of the regression models as more of the test data was used for
training, the relative order of performance among the regression models stayed constant in J. Deb
and coworkers’ experiments7. This signifies that a researcher can simulate these six for a smaller
percentage of the test data, say 10%, determine which model suits their need the best, and then
continue the training for that model only, saving time and energy.

2.3. Machine Learning Models Used for Data, Metadata, and


Algorithmic Data Extraction
Machine learning models such as large language models, support vector machines, and deep
neural networks can help build databases by extracting quantitative data, metadata, and
algorithmic data. The vast sea of literature and data floating online makes it overly time-
consuming or expensive to manually extract information by hand. For large-scale projects with
substantial budgets, extracting certain scientific data and metadata would be more plausible.
However, in the ceramics community, there hasn’t been enough effort or budget to create a
centralized and organized database hub for ceramics property data, as noted by S. Freiman and J.
Rumble4. As it stands, it is beneficial to be able to generate databases with high accuracy without
requiring extensive manpower, and a combination of machine learning models can help
researchers make those databases with minimal human input and without domain expertise.

2.3.1. Large Language Models for Quantitative Data Extraction


Despite the immensely practical and applicable nature of large language models (LLMs), it may
seem troublesome to use them to extract data, because with such a black-box system, one cannot
be entirely certain that the output is reliable. However, when combined with human supervision,
LLMs can extract data from research papers with impressive accuracy, as shown in the works of
M. Polak and coworkers5. The method of quantitative data extraction presented by M. Polak and
coworkers was able to achieve 90% precision at 96% recall and can generate databases of up to
around 1,000 entries within a single workday. Such a high precision and recall is noteworthy,
because even “state-of-the-art” named entity recognition-based tools such as the
ChemDataExtractor2 achieved a 52% precision at 37% recall in the same scenario 5. In M. Polak
and coworkers’ method, human input is necessary at the end of the process to manually extract
the data values from a list of sentences that were predicted to be relevant to the data in question.
However, “[fine-tuning] the LLM with example sentences” so that the LLM can more accurately
distinguish relevant sentences, significantly reduced the number of irrelevant sentences among
this mix, “about 99%” of them5. As the development of LLMs continues, using LLMs to extract
data from research papers seems to be a promising way to extract data without needing a data
scientist, as AI-powered LLMs such as ChatGPT can understand natural language as an input to
function.

2.3.2. Support Vector Machines for Metadata Extraction


Support vector machines (SVMs) are well known for their ability to accurately classify data
points of different classes. However, SVMs can be helpful in domains far beyond pure
classification tasks. One example is using SVMs to extract metadata from literature.

By dividing relevant metadata into fifteen classes, such as title, author, abstract, and more,
classifying each line as part of one or more of those classes, and iteratively correcting those
classifications by observing neighboring lines, H. Han and coworkers were able to extract
metadata fully automatically at accuracies above 98.5% for most of the fifteen classes of
metadata, an accuracy “nominally better” than other available methods of metadata extraction 15.

2.3.3. Deep Neural Networks for Algorithmic Data Extraction


To make the ceramics property databases adhere to the “interoperable” aspect of the FAIR
principle, it is important to provide access to the code used to analyze the data present or provide
a way to obtain such code without having to train domain experts in data science. Machine
learning models, specifically deep neural networks (DNNs), can be used to extract algorithmic
data so that the databases can contain pseudocode extracted alongside the data extracted using
large language models. I. Safder and coworkers were able to create a procedure dedicated to
extracting algorithmic data using clever techniques, showcased in Figure 616.
Figure 6: Depiction of Procedure for Extracting Algorithmic Data from Literature, Adapted from Deep Learning-
based Extraction of Algorithmic Metadata in Full-Text Scholarly Documents

As depicted in Figure 6, they first converted the PDF files into text files and looked for “sparse
boxes” that were left sparse because of the formatting required for showcasing programming
languages. Then, based on 60 features, including font, structure, and content of the text region,
the DNN was trained to classify whether the text given contains algorithmic data or not. The
classification resulting from this procedure had an F1-score of 93.32%, “outperforming the state-
of-the-art techniques by 28%”, and a near 80% accuracy16.
2.3.4. Flowchart of Machine Learning Methods to Utilize
Figure 7 depicts a flowchart to help guide which machine learning models to use and how to use
them to create databases through extraction.

Figure 7: Proposed Method of Database Creation Utilizing Machine Learning Models

First, from resources with an abundance of research papers, researchers can select candidates that
may be relevant to the material property in question. Then, using the metadata extraction method
with SVMs (2.3.2.), they can extract various metadata, including title, author, publication date,
and, most importantly, the abstract. If the researchers are aiming to create a large database, it
may be necessary to train an LLM to classify the documents as relevant to the property in
question or not, based on their abstracts, as it may be overwhelming to sift through all the
abstracts. However, for smaller databases, it may be more efficient to sift through summaries of
the abstracts manually to look for key phrases, especially if those summaries can be generated
quickly. According to D. Morgan, LLMs such as ChatGPT were “less successful at locating
subtle, interpretive themes, and more successful at reproducing concrete, descriptive themes” 17.
In other words, the LLM would be relatively reliable in summarizing and interpreting abstracts
of scientific papers, usually written in a concrete and factual tone. After selecting the batch of
relevant documents, the researchers can now use the quantitative data extraction method using
LLMs (2.3.1.) to obtain the numerical data in question. Finally, they can extract any algorithmic
data that can be used to analyze the data presented in various research papers, if any exists, to
implement as resources for the database extracted. Overall, the resulting database would have the
metadata required to decide whether the experiment or simulation used to obtain the data was
reliable, and the algorithmic data to interoperate the data.

3. Discussion
One of the biggest challenges the ceramic community must overcome for innovative research, as
presented by E. Guire and coworkers, is the severe lack of infrastructure and a data-sharing
platform. Currently, the ceramic community is given a minimal number of databases to research
with, and more databases that adhere to the FAIR principles and promote data-sharing are
necessary to accelerate data-driven research processes18. Using the proposal described in 2.3.4.,
we can create structured databases with metadata to validate the integrity of the data included.

Additionally, machine learning models have the potential to extrapolate data outside the bounds
of the data they were trained on. By training regression models such as the various decision trees,
a graph neural network, or support vector regression models, researchers can extrapolate
outwards from a set of training data to achieve predictions with incredible accuracy, following
the proposal described in 2.2.3.

Another challenge present in the ceramics society is the limited data accessibility and usability of
ceramic materials18. We believe that the methods for creating databases discussed in this paper
offer a reassuring prospect to help resolve this challenge. Through the methods proposed, large
organizations such as the American Ceramic Society can create a hub for a comprehensive list of
relevant ceramic materials data. This implies the possibility of a centralized database that one
can access with a single access point and subscription or signup, reducing the redundancy of
databases over the World Wide Web and improving the efficiency of researchers. Additionally,
the algorithmic data extracted using the methodology proposed in 2.3.4. allows ceramics
researchers to analyze and utilize the data in the databases to their fullest potential without
having to contact a domain expert in data science. This will reduce the gap between the ceramics
researchers and the data scientists, allowing the researchers to acquire the code necessary on the
spot and accelerating the research process. Further, if no algorithmic data were to be present, the
centralized database hub could implement a virtual assistant based on LLMs that can generate
the code needed for the user on demand, just like how ChatGPT can generate functioning code
snippets upon natural language instructions.

Another significant challenge the ceramic community must overcome is the lack of appropriate
computational approaches to accurately calculate large-scale models of ceramics18. Machine
learning is one of the keys to solving this challenge. By training machine learning models on
high-fidelity simulation data to build accurate and universal interatomic potentials, they can then
be used to simulate larger systems at a fraction of the computing time required by DFT. Some
commercial platforms have already broken ground in this area in catalysis, organic chemistry,
and polymers, but their application to ceramics remains limited. Considering how graph neural
networks can input and predict excited and ground state molecular structures of materials
accurately at a fraction of the time of a DFT simulation, further research into machine learning
and their applications to ceramics has the potential for scientific breakthroughs in the future.

Finally, we would like to caution the readers that the recommendations presented here are based
on a limited literature survey and are yet to be proven in practice. They should be taken as a first
step towards addressing challenges facing sustainability research in materials science in the field
of ceramics.

4. Methods
We have mainly used the Google Scholars search engine to search prompts that relate various
machine learning models (e.g., random forest decision tree), tasks (e.g., regression), or programs
(e.g., Scikit-learn) to ceramics or materials research to find research papers, journals, and articles
relevant to this paper. To verify the credibility of old sources, namely the journal written by S.
Freiman et al., we’ve done further research, manually visiting databases such as the National
Institute of Standards and Technology to validate whether the challenges mentioned in the
journal persist. We also evaluated the publication date, which is of the greatest importance for
the relevancy of machine learning methods for contemporary database creation, and how
relevant the proposed methods are for the purpose of creating ceramic property databases. To
assess the quality of the methods proposed, we’ve analyzed the design of the study, such as what
type of data (e.g. experimental, DFT calculated) was used to train the machine learning models
and the statistical measures provided. Also, we have visited websites created by reputable
sources, such as IBM, to obtain general information regarding machine learning models. To
organize the information obtained from surveying these sources, we’ve categorized these sources
based on the machine learning task performed and the model used, which aligns with the overall
structure of this paper as well.
5. References
1. European Ceramic Industry Association and others. Ceramic roadmap to 2050. Brussels,
Belgium: The European Ceramic Industry Association (2021).
2. K. Gong, K. Yang, C. E. White. Density functional modeling of the binding energies
between aluminosilicate oligomers and different metal cations. Frontiers in Materials 10
(2023).
3. S. Kim, M. Lee, C. Hong, Y. Yoon, H. An, D. Lee, W. Jeong, D. Yoo, Y. Kang, Y.
Youn, S. Han. A band-gap database for semiconducting inorganic materials calculated
with hybrid functional. Scientific Data 7, 387 (2020).
4. S. Freiman, J. Rumble. Current availability of ceramic property data and future
opportunities. American Ceramic Society Bulletin 92, 34–39 (2013).
5. M. Polak, S. Modi, A. Latosinska, J. Zhang, C. Wang, S. Wang, A. Hazra, D. Morgan.
Flexible, model-agnostic method for materials data extraction from text using general
purpose language models. Digital Discovery 3, 1221 (2023).
6. IBM. What are support vector machines (SVMs)? [Link]
vector-machine (2023).
7. J. Deb, J. Gou, H. Song, C. Maiti. Machine learning approaches for predicting the
ablation performance of ceramic matrix composites. Journal of Composites Science 8, 96
(2024).
8. IBM. What is a decision tree? [Link] (2023).
9. IBM. What is random forest? [Link] (2023).
10. IBM. What is a neural network? [Link] (2023).
11. IBM. What are large language models (LLMs)? [Link]
language-models (2023).
12. C. M. Clausen, J. Rossmeisi, Z. Ulissi. Adapting OC20-trained EquiformerV2 models for
high-entropy materials. [Link] (2024).
13. K. Kaufmann, D. Maryanovsky, W. Mellor, C. Zhu, A. Rosengarten, T. Harrington, C.
Oses, C. Toher, S. Curtarolo, K. Vecchio. Discovery of high-entropy ceramics via
machine learning. NPJ Computational Materials 6, 42 (2020).
14. F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel,
P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M.
Brucher, M. Perrot, E. Duchesnay, G. Louppe. Scikit-learn: machine learning in python.
Journal of Machine Learning Research 12, 2825–2830 (2011).
15. H. Han, C. Giles, E. Manavoglu, H. Zha, Z. Zhang, E. Fox. Automatic document
metadata extraction using support vector machines. 2003 Joint Conference on Digital
Libraries, 2003. Proceedings. 37–48 (2003).
16. I. Safder, S. Hassan, A. Visvizi, T. Noraset, R. Nawaz, S. Tuarob. Deep learning-based
extraction of algorithmic metadata in full-text scholarly documents. Information
Processing & Management 57, 102269 (2020).
17. D. Morgan. Exploring the use of artificial intelligence for qualitative data analysis: the
case of ChatGPT. International Journal of Qualitative Methods 22 (2023).
18. E. Guire, L. Bartolo, R. Brindle, R. Devanathan, E. Dickey, J. Fessler, R. French, U.
Fotheringham, M. Harmer, E. Lara-Curzio, S. Lichtner, E. Maillet, J. Mauro, M.
Mecklenborg, B. Meredi, K. Rajan, J. Rickman, S. Sinnott, C. Spahr, R. Weber. Data‐
driven glass/ceramic science research: Insights from the glass and ceramic and data
science/informatics communities. Journal of the American Ceramic Society 102 (2019).

You might also like