Example Standard References
Example Standard References
0. Abstract
Ceramics have been a witness of human ingenuity throughout history. From their usage in
pottery and artifacts in the past to industrial, electrical, and aerospace applications of ceramics in
the present, ceramics are ubiquitous in our lives. Unfortunately, manufacturing ceramics is
extremely energy-intensive and is responsible for more than 8% of global anthropogenic CO2
emissions. There have been continuous efforts to make these valuable materials more
sustainable. Recently, with the advance of data science, organizing and learning from data has
become much less time-consuming. This literature review aims to coalesce the two fields,
sustainability, and data science, with a focus on ceramics. Instead of narrowing our focus to
specific ceramics, such as glass or cement, we proposed procedures that can be broadly applied.
By analyzing different data science approaches and their applicability to creating ceramics
databases, we evaluated various ways to promote ceramics sustainability using data science.
More specifically, we analyzed vital insights or techniques that may be applied elsewhere and
sought opportunities to create procedures that required the least human input and benefited the
most from machine learning models. We categorized different approaches toward regression and
data extraction, examined suitable approaches for various scenarios, and qualified these efforts
by explaining the impact new databases can have on the sustainability of ceramics. We hope this
literature review will prove relevant to readers with an interest in data science, sustainable
ceramics, and their intersections.
Keywords: Ceramics, Sustainability, Data Science, Machine Learning (ML), Deep Learning,
Large Language Model (LLM), Regression, Data Extraction
1. Introduction
Ceramics are defined as materials that are neither metallic nor organic. These materials often
present themselves with valuable properties, such as heat resistance, semiconducting behavior, or
inertness. For example, refractories are used in extreme-temperature industrial processes because
of their high resistance to heat, and technical ceramics are used in chemical processing plants
because of their chemical inertness1. Ceramics are undoubtedly valuable, but unfortunately,
producing ceramics is a highly energy-intensive process. Aside from the mechanical energy
needed to mill and pulverize the mined raw materials (e.g., rocks of clays, silicas, limestones),
these raw materials require incredibly high temperatures (e.g., above 1500 °C in specially
equipped furnaces) to react in solid or molten form and yield the solid solutions that define the
products. The associated greenhouse gas emissions are not attributable only to the energy input
to the manufacturing process (e.g., process heat, electricity, mechanical power) but also to the
process emissions emanating from the solid-state reactions responsible for chemically converting
the raw materials to products, accentuating the climate impact of the ceramics industry via these
process emissions. For instance, emissions from the European ceramic industry amount to 19
million tons of CO2 annually1, and much research is necessary to make the ceramics industry
more sustainable.
We provide a general schematic of the ceramics industry below to better explain the
sustainability challenge in ceramics. After mining from Earth deposits, raw materials such as
clay, silica, or limestone are transported to and stored in the manufacturing plant. This
transportation process could lead to significant emissions if raw materials are mined far from the
ceramic plant. The raw materials at the ceramic plant are then milled, mixed, and pulverized to
increase the surface area available for the ensuing solid-state reactions. Often, they get formed
into a paste and shaped into a mold before getting dried, to prepare them for the firing step. The
raw materials then react under firing in furnaces, where heat management and distribution are of
paramount importance to the quality of the product. Afterward, finishing often includes cooling,
packaging, and regular testing for quality control of the products.
Figure 1: A Simplified General Schematic of the Ceramics Manufacturing Process adapted from Ceramic Roadmap
to 2050
Figure 1 overlays some examples of opportunities for sustainability in existing ceramic plants.
Following the schematic above, one can start by placing the ceramics manufacturing site near the
mining and raw materials extraction site to minimize transportation emissions. The drying step is
crucial before firing, as this step controllably gets rid of water and volatile organic impurities
contaminating the raw materials, ensuring the mechanical integrity of the fragile and, as of yet,
wet shapes. Here, there exists an opportunity for heat integration, where the flue gases from the
firing step heat the drying air or are applied directly to dry the incoming feed to the furnace.
Using environmentally sustainable fuel sources such as green hydrogen during this firing stage
could significantly improve the sustainability of the overall process. The furnace is also lined
with ceramic refractories that trap the heat inside, offering effective insulation. Further, capturing
the carbon released during the drying and firing stages can help decarbonize the production
process. However, further research must be done to find more practical carbon removal
technologies and offsetting measures because current methods, like carbon capture and storage,
are too costly to be practical. The ceramics industry is heading towards a circular pathway as
well, in which refuse from manufacturing and distribution sites is used as a filler to minimize
new raw materials in the production mix, solving both a waste problem and an emissions
problem. Some wall and floor tiles industry manufacturers transitioned from rotary printing to
digital printing and created walls and floor tiles with 80% recycled ceramics using ceramic inks
instead of decorative pastes. Clay pipes consisting of 40% new raw materials and refractories
containing 20% to 80% new raw materials can be produced, and it is possible to reuse 90% of
expanded clay products1. The goal should be to make all ceramic products 100% reusable so no
waste is created. Although this might require significant research or technological breakthroughs,
decreasing the percentage of raw materials used would result in substantial reductions in process
emissions.
However, there is only so much that such improvements can do to turn the ceramics industry into
a genuine engine for sustainability, for the ceramics industry remains intrinsically emissions-
intensive and geared towards end products rather than advanced ceramics that go into further
green technologies1. Here, we suggest that a smart approach towards sustainable ceramics
research is data-centered, including building, maintaining, and seamlessly interacting with
comprehensive databases that adhere to FAIR data principles (findable, accessible, interoperable,
and reusable).
To make this point, we highlight two recent articles discussing the highly successful results of
this endeavor. The first is by K. Gong and coworkers at Princeton University, where they
established a strong research program geared towards the manufacture of cement with less
emissions, by designing alkali-activated kaolin clays as raw materials that help turn industrial
wastes and already-calcined clays (which emit no further CO2 upon processing) into cementitious
binders, minimizing process emissions2. Specifically, they applied computationally intensive
density functional theory simulations to calculate the binding energies of seventeen such alkali
cations to aluminosilicate dimers and trimers and then correlated these binding energies to two
other properties that are readily obtainable (i.e., do not require DFT simulations): the ionic
potential and field strength of these cations. The high-quality second-order polynomial
correlations (R2=0.99-1.00) highlight their usability to rapidly estimate such binding energies of
new cations or new substrates, extending this research's impact into materials not initially
simulated by the authors2. The data on binding energies of different cations to different clay
materials can help design alkali-activated kaolin clays as raw materials by providing insight into
the performances (e.g., performance against sulfuric acid attacks) of the resulting sustainable
cement. This endeavor highlights the importance of a data-centered approach toward innovative
ceramics research.
The second article is by S. Kim and coworkers at Seoul National University. They successfully
created an extensive band-gap database for semiconducting inorganic materials, as applied in
photovoltaic devices. By using a hybrid functional and considering a stable magnetic ordering,
they made a database for band gaps for 10,841 materials with a significantly smaller root-mean-
square error of 0.36 eV than the 0.75-1.05 eV of existing databases (1 eV ~ 100 kJ/mol)3.
Additionally, they identified a considerable number of small-gap materials that were
misclassified as metals in other databases. By correctly classifying these materials as
semiconducting inorganics, they provided a broader selection of materials that manufacturers or
researchers may consider using. This can come with various benefits, such as reduced costs of
materials, reduced transport emissions, and more. Overall, S. Kim and coworkers created a
highly impactful (and highly cited) database utilizing a data-centered approach, further
highlighting the potential of a data-centered approach for scientific research and innovation.
Band gaps are but one of a myriad of properties that define the utility of ceramic materials in
advanced and sustainable applications (see Table 1 for a non-comprehensive list of such
properties and applications).
Table 1: A Simplified Subset of Major Property Categories for Technical and Sustainable Ceramics
Oxygen Ionic Ability to Solid oxide fuel Electrical S/m (siemens per
Conductivity conduct oxygen cells, oxygen meter)
ions under an sensors
applied voltage
gradient
2. Results
2.1. An Overview of Machine Learning Models
Following the recent rise of machine learning, there have been recent efforts to extract data from
text using natural language processing models. For example, M. Polak and coworkers proposed a
successful procedure using ChatGPT to extract mid-sized databases from text5. This process
requires minimal prior knowledge of coding and the property being analyzed, but can generate
databases with high precision and recall in a single day. This automatic quality of machine
learning, being able to extract data from literature or learn from available data to create
sufficiently accurate databases with minimal human input, is what makes machine learning so
valuable and, thus, will be the focus of this literature review.
In the world of data science, machine learning methods have varying types of training data and
tasks to solve. Supervised models or algorithms refer to models trained on data sets with labels
— either categorical or continuous — corresponding to the data points, and can help predict
material properties. Some supervised models, such as graph neural networks, can input graphs as
data, which is useful when predicting properties based on molecular structures that are
represented as graphs. In contrast, unsupervised models are trained on data sets that lack such
labels and are often used for clustering or anomaly detections, but lack utility for the purpose of
creating databases for ceramics data. This paper will focus on the following three most common
types of supervised machine learning models that were present in the literature surveyed:
Support Vector Machines, Decision Trees, and Neural Networks. A brief overview of these three
types of models is described in Figure 2.
Figure 2: A Simple Overview of Three Types of Machine Learning Models
A random forest decision tree is an ensemble of decision trees trained separately on random
samples of the training data. These independent decision trees are run in parallel when given an
input, and the individual results generated by each of these decision trees are aggregated to make
a final prediction. A gradient boost decision tree is a series of several decision trees that attempt
to predict the errors from the previous decision trees in the series. An extreme gradient boost
decision tree is a gradient boost decision tree with regularizations that prevent complex models
that are overfitting the training data and/or aren’t generalizing well. These regularizations
include but are not limited to, lasso regularizations, which encourage sparsity, and ridge
regularizations, which reduce the variance of the model. An adaptive boost decision tree consists
of a group of independent decision trees that learn adaptively, forcing decision trees that made
significant errors to train more on those instances to reduce the large errors. The results from
these independent decision trees are then weighted and compiled so that the decision trees that
made those significant errors have less influence over the result7 8 9.
Graph Neural Networks are a type of neural network that can intake graph-structured irregular
data. These neural networks are especially useful in understanding materials because they can
input molecular structures as data and handle them efficiently. Deep Neural Networks,
commonly referred to as Deep Learning, are neural networks with three or more layers, which is
common for models that require complexity. These models can be trained to estimate material
properties using other more available properties without knowing the specific relation between
those properties beforehand. A Large Language Model (LLM) is an application of deep learning
that utilizes multiple layers of deep neural networks for natural language processing and is
trained on extremely large samples of textual data. These various types of neural networks can be
useful in ceramics research because of their ability to interpret large amounts of data or text, to
predict numerical data, or to extract different data from the vast sea of literature 11.
2.2. Machine Learning Models Used for Regression
Machine learning models such as deep neural networks and the various decision trees can help
solve regression tasks by using data that are more easily accessible to predict values of properties
that are expensive or computationally intensive to compute. This is a common practice for
scientific researchers, as seen by how K. Gong and coworkers successfully calculated binding
energies using density functional theory (DFT) and then extrapolated them with second-order
polynomials to avoid further DFT calculations on metal cations not previously calculated 2. The
main benefit of using machine learning models to make these predictions is that they can
calculate predictions nearly instantaneously with no human input or background information on
these properties.
Unfortunately, some decision tree algorithms are much more complicated in nature and are
harder to interpret than random forest decision trees or regular decision trees. Such is the case for
extreme gradient boost and adaptive boost decision trees. Because normalizations and
optimizations are applied to extreme gradient boost trees and because adaptive boost decision
trees compare the training data, they are intrinsically too complex to interpret or understand.
Still, these decision trees take comparatively much less time to make predictions. For example,
when J. Deb and coworkers predicted the ablation performance of ceramic matrix composites
using various decision trees to compare results, the extreme gradient boost and adaptive boost
decision trees took significantly less simulation time7. Specifically, when using the performance
parameters that resulted in the greatest R-squared scores and the least mean absolute errors
among all decision trees, the extreme gradient boost and adaptive boost decision trees took 1
second and 2 seconds accordingly, the random forest decision tree took slightly longer than 2
minutes, the regular decision tree took longer than 18 minutes, and the gradient boost decision
tree took close to 9 minutes7. This shows that the decision trees that are too complicated to be
interpreted still serve a purpose, as they can drastically improve computing efficiency while still
retaining high accuracy.
2.2.3. Flowchart of Machine Learning Methods to Utilize for Different
Regression Tasks
Figure 5: A Brief Overview of What Machine Learning Models to Use Given Different Scenarios
The flowchart depicted in Figure 5 shows a summary of what machine learning models one
should use depending on the type of regression task assigned. The properties of ceramics are
dependent on the molecular structure of the material. As such, it is plausible that some form of
molecular structure data is involved with the inputs or outputs of a given regression task. In such
cases, it is encouraged to use a deep neural network to be able to take full advantage of the
structured data. Further, predicting molecular structures of materials as an intermediate step
before predicting a quantitative property could be useful, as knowing the molecular structures
can help determine the distribution of electron density across the material, find distances between
atoms, and find angles between bonds, all of which allow us to utilize DFT or quantum
mechanical modeling to predict other properties of a material, if such level of accuracy is
necessary for a project or experiment. For example, a graph neural network can be used to help
narrow down which candidate has the significance to further test using expensive DFT
calculations to make more accurate results.
When predicting continuous, quantitative data, you can utilize various decision trees or a support
vector regression. In particular, the extreme gradient boost decision tree, adaptive boost decision
tree, and random forest decision tree models execute the fastest among these six machine
learning models shown in Figure 5. This was shown experimentally in J. Deb and coworkers’
experiments7, and aligns with the properties of the model: the extreme gradient boost tree is an
optimization of a gradient boost tree that implements parallelization to speed up the tree-building
process, while the adaptive boost decision tree and the random forest decision trees generate
small trees to aggregate their results, which takes less time than generating a larger, more
complex tree. Other than the apparent difference in the expected execution time of these models,
it is hard to predict which model will perform the best in terms of accuracy or precision because
of how unique these models are. Here, it is worth noting that the comparative performance of
these models remains relatively constant even as the test data increases. Despite the small
changes in the performances of the regression models as more of the test data was used for
training, the relative order of performance among the regression models stayed constant in J. Deb
and coworkers’ experiments7. This signifies that a researcher can simulate these six for a smaller
percentage of the test data, say 10%, determine which model suits their need the best, and then
continue the training for that model only, saving time and energy.
By dividing relevant metadata into fifteen classes, such as title, author, abstract, and more,
classifying each line as part of one or more of those classes, and iteratively correcting those
classifications by observing neighboring lines, H. Han and coworkers were able to extract
metadata fully automatically at accuracies above 98.5% for most of the fifteen classes of
metadata, an accuracy “nominally better” than other available methods of metadata extraction 15.
As depicted in Figure 6, they first converted the PDF files into text files and looked for “sparse
boxes” that were left sparse because of the formatting required for showcasing programming
languages. Then, based on 60 features, including font, structure, and content of the text region,
the DNN was trained to classify whether the text given contains algorithmic data or not. The
classification resulting from this procedure had an F1-score of 93.32%, “outperforming the state-
of-the-art techniques by 28%”, and a near 80% accuracy16.
2.3.4. Flowchart of Machine Learning Methods to Utilize
Figure 7 depicts a flowchart to help guide which machine learning models to use and how to use
them to create databases through extraction.
First, from resources with an abundance of research papers, researchers can select candidates that
may be relevant to the material property in question. Then, using the metadata extraction method
with SVMs (2.3.2.), they can extract various metadata, including title, author, publication date,
and, most importantly, the abstract. If the researchers are aiming to create a large database, it
may be necessary to train an LLM to classify the documents as relevant to the property in
question or not, based on their abstracts, as it may be overwhelming to sift through all the
abstracts. However, for smaller databases, it may be more efficient to sift through summaries of
the abstracts manually to look for key phrases, especially if those summaries can be generated
quickly. According to D. Morgan, LLMs such as ChatGPT were “less successful at locating
subtle, interpretive themes, and more successful at reproducing concrete, descriptive themes” 17.
In other words, the LLM would be relatively reliable in summarizing and interpreting abstracts
of scientific papers, usually written in a concrete and factual tone. After selecting the batch of
relevant documents, the researchers can now use the quantitative data extraction method using
LLMs (2.3.1.) to obtain the numerical data in question. Finally, they can extract any algorithmic
data that can be used to analyze the data presented in various research papers, if any exists, to
implement as resources for the database extracted. Overall, the resulting database would have the
metadata required to decide whether the experiment or simulation used to obtain the data was
reliable, and the algorithmic data to interoperate the data.
3. Discussion
One of the biggest challenges the ceramic community must overcome for innovative research, as
presented by E. Guire and coworkers, is the severe lack of infrastructure and a data-sharing
platform. Currently, the ceramic community is given a minimal number of databases to research
with, and more databases that adhere to the FAIR principles and promote data-sharing are
necessary to accelerate data-driven research processes18. Using the proposal described in 2.3.4.,
we can create structured databases with metadata to validate the integrity of the data included.
Additionally, machine learning models have the potential to extrapolate data outside the bounds
of the data they were trained on. By training regression models such as the various decision trees,
a graph neural network, or support vector regression models, researchers can extrapolate
outwards from a set of training data to achieve predictions with incredible accuracy, following
the proposal described in 2.2.3.
Another challenge present in the ceramics society is the limited data accessibility and usability of
ceramic materials18. We believe that the methods for creating databases discussed in this paper
offer a reassuring prospect to help resolve this challenge. Through the methods proposed, large
organizations such as the American Ceramic Society can create a hub for a comprehensive list of
relevant ceramic materials data. This implies the possibility of a centralized database that one
can access with a single access point and subscription or signup, reducing the redundancy of
databases over the World Wide Web and improving the efficiency of researchers. Additionally,
the algorithmic data extracted using the methodology proposed in 2.3.4. allows ceramics
researchers to analyze and utilize the data in the databases to their fullest potential without
having to contact a domain expert in data science. This will reduce the gap between the ceramics
researchers and the data scientists, allowing the researchers to acquire the code necessary on the
spot and accelerating the research process. Further, if no algorithmic data were to be present, the
centralized database hub could implement a virtual assistant based on LLMs that can generate
the code needed for the user on demand, just like how ChatGPT can generate functioning code
snippets upon natural language instructions.
Another significant challenge the ceramic community must overcome is the lack of appropriate
computational approaches to accurately calculate large-scale models of ceramics18. Machine
learning is one of the keys to solving this challenge. By training machine learning models on
high-fidelity simulation data to build accurate and universal interatomic potentials, they can then
be used to simulate larger systems at a fraction of the computing time required by DFT. Some
commercial platforms have already broken ground in this area in catalysis, organic chemistry,
and polymers, but their application to ceramics remains limited. Considering how graph neural
networks can input and predict excited and ground state molecular structures of materials
accurately at a fraction of the time of a DFT simulation, further research into machine learning
and their applications to ceramics has the potential for scientific breakthroughs in the future.
Finally, we would like to caution the readers that the recommendations presented here are based
on a limited literature survey and are yet to be proven in practice. They should be taken as a first
step towards addressing challenges facing sustainability research in materials science in the field
of ceramics.
4. Methods
We have mainly used the Google Scholars search engine to search prompts that relate various
machine learning models (e.g., random forest decision tree), tasks (e.g., regression), or programs
(e.g., Scikit-learn) to ceramics or materials research to find research papers, journals, and articles
relevant to this paper. To verify the credibility of old sources, namely the journal written by S.
Freiman et al., we’ve done further research, manually visiting databases such as the National
Institute of Standards and Technology to validate whether the challenges mentioned in the
journal persist. We also evaluated the publication date, which is of the greatest importance for
the relevancy of machine learning methods for contemporary database creation, and how
relevant the proposed methods are for the purpose of creating ceramic property databases. To
assess the quality of the methods proposed, we’ve analyzed the design of the study, such as what
type of data (e.g. experimental, DFT calculated) was used to train the machine learning models
and the statistical measures provided. Also, we have visited websites created by reputable
sources, such as IBM, to obtain general information regarding machine learning models. To
organize the information obtained from surveying these sources, we’ve categorized these sources
based on the machine learning task performed and the model used, which aligns with the overall
structure of this paper as well.
5. References
1. European Ceramic Industry Association and others. Ceramic roadmap to 2050. Brussels,
Belgium: The European Ceramic Industry Association (2021).
2. K. Gong, K. Yang, C. E. White. Density functional modeling of the binding energies
between aluminosilicate oligomers and different metal cations. Frontiers in Materials 10
(2023).
3. S. Kim, M. Lee, C. Hong, Y. Yoon, H. An, D. Lee, W. Jeong, D. Yoo, Y. Kang, Y.
Youn, S. Han. A band-gap database for semiconducting inorganic materials calculated
with hybrid functional. Scientific Data 7, 387 (2020).
4. S. Freiman, J. Rumble. Current availability of ceramic property data and future
opportunities. American Ceramic Society Bulletin 92, 34–39 (2013).
5. M. Polak, S. Modi, A. Latosinska, J. Zhang, C. Wang, S. Wang, A. Hazra, D. Morgan.
Flexible, model-agnostic method for materials data extraction from text using general
purpose language models. Digital Discovery 3, 1221 (2023).
6. IBM. What are support vector machines (SVMs)? [Link]
vector-machine (2023).
7. J. Deb, J. Gou, H. Song, C. Maiti. Machine learning approaches for predicting the
ablation performance of ceramic matrix composites. Journal of Composites Science 8, 96
(2024).
8. IBM. What is a decision tree? [Link] (2023).
9. IBM. What is random forest? [Link] (2023).
10. IBM. What is a neural network? [Link] (2023).
11. IBM. What are large language models (LLMs)? [Link]
language-models (2023).
12. C. M. Clausen, J. Rossmeisi, Z. Ulissi. Adapting OC20-trained EquiformerV2 models for
high-entropy materials. [Link] (2024).
13. K. Kaufmann, D. Maryanovsky, W. Mellor, C. Zhu, A. Rosengarten, T. Harrington, C.
Oses, C. Toher, S. Curtarolo, K. Vecchio. Discovery of high-entropy ceramics via
machine learning. NPJ Computational Materials 6, 42 (2020).
14. F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel,
P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M.
Brucher, M. Perrot, E. Duchesnay, G. Louppe. Scikit-learn: machine learning in python.
Journal of Machine Learning Research 12, 2825–2830 (2011).
15. H. Han, C. Giles, E. Manavoglu, H. Zha, Z. Zhang, E. Fox. Automatic document
metadata extraction using support vector machines. 2003 Joint Conference on Digital
Libraries, 2003. Proceedings. 37–48 (2003).
16. I. Safder, S. Hassan, A. Visvizi, T. Noraset, R. Nawaz, S. Tuarob. Deep learning-based
extraction of algorithmic metadata in full-text scholarly documents. Information
Processing & Management 57, 102269 (2020).
17. D. Morgan. Exploring the use of artificial intelligence for qualitative data analysis: the
case of ChatGPT. International Journal of Qualitative Methods 22 (2023).
18. E. Guire, L. Bartolo, R. Brindle, R. Devanathan, E. Dickey, J. Fessler, R. French, U.
Fotheringham, M. Harmer, E. Lara-Curzio, S. Lichtner, E. Maillet, J. Mauro, M.
Mecklenborg, B. Meredi, K. Rajan, J. Rickman, S. Sinnott, C. Spahr, R. Weber. Data‐
driven glass/ceramic science research: Insights from the glass and ceramic and data
science/informatics communities. Journal of the American Ceramic Society 102 (2019).