DiffBatt: Predicting Battery Degradation
DiffBatt: Predicting Battery Degradation
1
Institute for Software and Systems Engineering,
Clausthal University of Technology, Germany
2
Institute of Chemical and Electrochemical Process Engineering,
Clausthal University of Technology, Germany
3
Institute of Automotive Management and Industrial Production,
Technische Universität Braunschweig, Germany
4
Institute of Machine Tools and Production Technology,
Technische Universität Braunschweig, Germany
5
Battery LabFactory Braunschweig (BLB),
Technische Universität Braunschweig, Germany
Abstract
1 Introduction
1.1 Lithium-ion batteries
Lithium-ion (Li-ion) batteries are key technologies in the field of energy storage, with applications
spanning portable electronics and electric vehicles [4]. The prominence of these batteries is largely
attributable to their high energy density, which enables substantial energy storage within a compact
and lightweight form factor. Moreover, Li-ion batteries demonstrate an extended cycle life compared
∗
Corresponding author: he76@[Link] ([Link]
Foundation Models for Science Workshop,38th Conference on Neural Information Processing Systems (NeurIPS
2024).
to other battery technologies, quantified in terms of charge and discharge cycles, thereby enhancing
their cost-effectiveness for long-term usage. The low self-discharge rate of Li-ion batteries ensures
minimal energy loss during periods of inactivity, which is a significant advantage over other battery
technologies [16, 41]. Nevertheless, several challenges remain. Safety continues to be a major issue,
as mechanical damage or improper handling can potentially lead to hazardous events such as thermal
runaway [8]. Furthermore, the economic and environmental implications of Li-ion battery production,
recycling, and disposal present additional complexities that warrant ongoing investigation [3, 17].
One persistent challenge that continues to impact their long-term performance and reliability is
capacity degradation.
The phenomenon of capacity degradation in Li-ion batteries is a multifaceted issue that encompasses
both the effects of aging and the effects of cycling. Aging behavior, which is often referred to as
calendar aging, pertains to the decline in battery performance over time, irrespective of active usage.
Factors such as ambient temperature, state of charge, and storage conditions play a significant role
in this degradation mode. In contrast, cycling behavior, also termed cycle aging, is linked to the
deterioration that batteries experience during charge and discharge cycles. High charge-discharge
rates and frequent cycling result in the accumulation of irreversible changes within the battery’s
electrochemical structure. This degradation is driven by a number of factors, including the formation
of a solid electrolyte interphase layer, electrolyte decomposition, and the growth of lithium plating
[7, 13]. Both aging and cycling behaviors collectively result in overall degradation, reducing the
battery’s capability to store and deliver electric charge [39] over its operational lifespan. In addition,
the aging of batteries is very individual depending on, e.g., usage behavior and environmental
conditions, which makes a basic understanding difficult and hinders individual battery management.
Despite extensive research, accurately predicting the rate and extent of capacity loss remains a
formidable challenge [35]. Therefore, advanced modeling techniques, including machine learning,
are increasingly being employed to provide more accurate predictions of the battery degradation
processes.
The degradation of a battery can be quantified by using key performance indicators, including the
state of health (SOH) and the remaining useful life (RUL) [28]. The SOH is a measure of the current
capacity of a battery relative to its original capacity, expressed as a percentage, and provides insight
into the extent of capacity degradation [10, 38]. In contrast, the RUL is a predictive measure that
estimates the remaining operational cycles of a battery before it reaches defined performance criteria
[31].
Methods for battery degradation modeling can be classified according to Rauf et al. [37] into four
domains: i) physics-based models, ii) empirical models, iii) data-driven methods (DDMs), and iv)
hybrid methods. Among these various domains, DDMs are emerging as a prominent technique for
developing battery degradation models. This is due to the flexibility and independence from specific
model assumptions that these approaches offer. In the domain of DDMs, machine learning (ML)
methods are widely regarded as one of the most effective approaches for estimating RUL and SOH,
due to their ability to address non-linear problems [37]. Since all battery RUL and SOH prediction
tasks are effectively regression problems, supervised learning is the most commonly used approach
in ML battery studies.
Recent literature reviews [28, 31, 37, 38, 43] indicate that various ML methods are utilized for
modeling battery degradation. In the area of artificial neural networks, shallow neural networks can
capture nonlinear relationships among an arbitrary number of inputs and outputs, however, they are
hindered by slow training processes and a propensity to converge at local minima [28, 38]. In contrast,
deep learning algorithms demonstrate superior performance in managing large datasets due to their
specialized architectures. They provide higher accuracy and enhanced generalization capabilities but
incur significant computational costs [38]. Techniques such as convolutional neural networks (CNNs),
recurrent neural networks (RNNs), and long short-term memory (LSTM) networks are commonly
employed in this context [31, 37, 43]. Additionally, support vector machines achieve a commendable
balance between generalization capability and estimation accuracy. However, they may struggle
with scalability on larger datasets [38]. Similarly, relevance vector machines have the disadvantage
2
of requiring extensive datasets, which results in significant computational complexity. However,
they offer the advantage of high accuracy, robust learning capabilities, and the capacity to generate
predictions with associated probability distributions [37]. Lastly, Gaussian process regression (GPR)
methods are advantageous for their ability to quantify the uncertainty of estimated values, which is
particularly valuable in practical applications. Nonetheless, GPR methods typically exhibit lower
efficiency in high-dimensional spaces and can be computationally complex [28, 38]. Another recent
approach, with a sole focus on the prediction of the SOH, is presented by Luo et al. [32], in which the
authors introduce the methodology of diffusion models as a promising avenue for SOH prediction.
Despite the widespread application of ML models in battery degradation analysis, comparing these
various approaches presents significant challenges. Many studies utilize different datasets, which are
often not publicly available due to confidentiality concerns. To address this issue, BatteryML was
developed by Zhang et al. [47], offering a standardized method for data representation that consoli-
dates and harmonizes all accessible public battery datasets. Additionally, BatteryML establishes clear
benchmarks for predicting RUL and includes a range of models, such as linear models, tree-based
models, and neural networks, tailored for battery degradation prediction. Lastly, BatteryML also
introduces the transformer architecture as a novel approach for predicting SOH and RUL [47].
Our contribution represents a threefold advancement in the domain of battery degradation predic-
tion. First, we introduce DiffBatt, a novel application of denoising diffusion probabilistic models
(DDPMs) specifically tailored for estimating the SOH and RUL of Li-ion batteries. The presented
approach is motivated by the need to handle the stochastic and intricate nature of battery degradation.
This innovative approach employs the capabilities of diffusion models to more effectively capture
the complex degradation behavior than traditional methods (Section 2.2). Secondly, our results
demonstrate a significant improvement over existing benchmarks, showcasing superior accuracy and
reliability in predictions (Section 3). Finally, we establish the groundwork for further development
by positioning our model as a foundation model, enabling future enhancements and adaptations to
diverse energy storage related applications (Section 4). All the codes and pre-trained models are
available on GitHub: [Link]
2 Methodology
DDPMs [21] are an expressive and flexible family of generative models that utilize a parameterized
Markov chain to produce high-quality samples that match the characteristics of the training data.
The key idea behind DDPMs is to learn a reverse process, also known as the reverse diffusion or
denoising process, which gradually removes noise from a sequence of noisy samples until it reveals
the original signal. The forward process of a DDPM is a Markov chain that progressively adds noise
to the data in the opposite direction of sampling. This process continues until the signal is completely
obscured by noise. The model’s objective is to learn a set of transformations that can effectively undo
this noise accumulation and recover the original signal.
The training of DDPMs involves using variational inference to optimize the model parameters, with
the goal of minimizing the difference between the generated samples and the true data distribution.
This is achieved by estimating the lower bound of the loss function along a large number of diffusion
steps, which are computed iteratively during the training process. The resulting model can generate
new, diverse samples that resemble the original data, making it useful for applications such as image
and audio synthesis [12, 22, 45], data augmentation [33], and more [2, 14, 29, 46].
We follow the principles proposed by Ho and Salimans [20] for classifier-free diffusion guidance
to increase sample quality while decreasing sample diversity in diffusion models. Classifier-free
guidance serves to achieve similar objectives for performing truncated or low-temperature sampling
in certain generative models, such as generative adversarial networks (GANs) and flow-based models.
The intended outcome is a decrease in sample diversity accompanied by an increase in individual
sample quality. Examples are truncation in BigGAN [6] and low-temperature sampling in Glow [26],
which lead to a trade-off curve between the Fréchet inception distance (FID) score and the inception
score. This enables the flexibility to generate high-quality or more diverse samples when predicting
or synthesizing battery degradation, respectively.
3
DiffBatt is based on a DDPM enhanced with transformer models and utilizes classifier-free diffusion
guidance for conditional generative modeling. In the following, we provide a brief discussion on the
methodological aspects of DDPMs and classifier-free guidance.
2.1 Background
Denoising diffusion models learn to systematically transform a sample of a simple prior, typically
a unit Gaussian, to a sample from an unknown data distribution q(x). In a fixed forward process,
a given data sample x0 ∼ q(x) is corrupted by Gaussian noise according to a variance schedule
{βt ∈ (0, 1)}Tt=1 over the course of T timesteps
T
Y p
q(x1:T |x0 ) = q(xt |xt−1 ), q(xt |xt−1 ) = N (xt ; 1 − βt xt−1 , βt I). (1)
t=1
in which the unknown true inverse conditional distribution q(xt−1 |xt ) is approximated by a neural
network pθ (xt−1 |xt ) parameterized by θ. The model seeks to learn an estimator for the mean
parameter µθ (xt , t), under the constraint that the covariance remains unchanged as
1 − ᾱt−1
Σ (xt , t) = βt I = Σt I (3)
1 − ᾱt
Qt
with ᾱt = αi , αt = 1 − βt . The mean µθ (xt , t) is parameterized as
i=1
1 βt
µθ (xt , t) = √ xt − √ ϵθ (xt , t) , (4)
αt 1 − ᾱt
√ √
where xt = ᾱt x0 + 1 − ᾱt ϵt for ϵt ∼ N (0, I), ϵθ is a function approximator intended to predict
ϵt from xt and ϵt indicates Gaussian noise to diffuse x0 to xt . The reverse process is trained to
approximate the joint distribution of the forward process by optimizing the evidence lower bound.
With the parameterizations suggested in [22] the loss simplifies to
h i
2
Lsimple (θ) := Et,x0 ,ϵ ∥ϵt − ϵθ (xt , t)∥ , (5)
resembling denoising score matching over multiple noise scales. To summarize, we can train the
reverse process to predict ϵt . To learn a conditional model pθ (x0 |c) the diffusion model is extended
by incorporating the conditioning variable c into the reverse process
T
Y
pθ (x0:T |c) = p(xT ) q(xt−1 |xt , c),
t=1
(6)
pθ (xt−1 |xt , c) = N (xt−1 ; µθ (xt , t, c), Σθ (xt , t, c)).
Following classifier-free guidance [20], we train an unconditional DDPM alongside the conditional
one by randomly setting the conditioning information c to a null token with probability puncond ,
set as a hyperparameter. To generate samples, we combine the scores from the conditional and
unconditional models
ϵ̃θ (xt , t, c) = (1 + w)ϵθ (xt , t, c) − wϵθ (xt , t), (7)
where w is the guidance strength.
2.2 Architecture
Similar to the current state-of-the-art architectures for image and audio diffusion models [9, 45],
DiffBatt is based on a U-Net architecture (see Fig. 1a) and employs diffusion processes to generate
SOH curves similar to a time series generation task. Conditioning for battery information, e.g.,
the capacity matrix, or a diffusion timestep t, is provided by adding embeddings into intermediate
4
layers of the network [21]. The model consists of residual blocks with one-dimensional convolutions,
attention modules, and pooling and up-sampling layers. In the reverse process, the model takes a
one-dimensional Gaussian noise as input and generates a sample of SOH degradation. Notably, SOH
typically exhibits a decreasing trend with progressive cycling. To better capture this physical behavior,
we append a positional encoding to the output of the first convolution, allowing the denoising process
to incorporate knowledge of the cycle number. Figure 1 depicts a schematic view of the model
architecture.
a) Training loss
U-Net
+
SOH
after
k steps
Cycle number
Figure 1: Schematic view of the model architecture. Adapted and modified from the work by Fürrutter
et al. [14], with permission from the authors. Modifications include context-specific changes.
DiffBatt can transform Gaussian noise into a new SOH curve through the reverse diffusion process,
as illustrated in Fig. 1c and Fig. 2. For this study, we employ the concept of the capacity matrix (Q),
as introduced by Attia et al. [1], as an additional condition for the diffusion process. The capacity
matrix serves as a compact representation of battery electrochemical cycling data, incorporating a
series of feature representations. Consistent with prior research on machine learning for predicting
battery degradation [1, 40, 47], we utilize the capacity matrix corresponding to the first 100 cycles.
This choice is driven by the high costs, time, and effort associated with long-term battery testing. Our
goal is to leverage early life performance data to predict battery degradation and minimize resource
expenditure. To encode Q into an embedding (cq ), we utilize a transformer encoder (see Fig. 1b).
This allows DiffBatt to generate SOH curves, from which the RUL can be derived by calculating
the number of cycles until the SOH drops below a specified threshold, such as 80% of the nominal
capacity.
t=0 t = 100 t = 200 t = 300 t = 400 t = 500 t = 600 t = 700 t = 800 t = 900 t = 1000
SOH(%)
Figure 2: Denoising steps for one test sample of the MATR dataset.
2.3 Data
5
utilize the data splits provided by BatteryML to maintain consistency and comparability, and we
benchmark our results against those reported in their study. We refer to Appendix A.1 and Zhang
et al. [47] for more detailed information on the data.
We train DiffBatt with puncond of 0.2 on different datasets and compare it with the models from
BatteryML [47]. For the prediction tasks, we generate ten SOH samples from ten different input
noises for each sample of the capacity matrix. Then, we select the sample that best fits the first 100
cycles as the final prediction. The RUL is further computed from the predicted SOH curve.
The results for the RUL prediction task are summarized in Table 1. This table illustrates the
performance of the models in the RUL task using root-mean-squared error (RMSE) as the evaluation
metric. RMSE is suitable for the RUL task since it represents the error based on an average number
of cycles in which the predicted RUL differs from the reference. We train the DiffBatt model with ten
different initialization seeds for each test and report the mean error along with the standard deviation,
indicated as a subscript. DiffBatt shows notable performance improvements on multiple datasets.
Specifically, DiffBatt achieved the lowest RMSE on the MATR1, SNL, and CRUSH datasets, with RMSE
values of 88 ± 4, 125 ± 11, and 294 ± 18 respectively. These results highlight DiffBatt’s robustness
and precision in predicting RUL across different battery compositions and operational conditions. For
instance, in the SNL dataset, DiffBatt’s RMSE of 125 ± 11 outperforms the best benchmark model,
PCR, which has an RMSE of 200 and is significantly lower than that of the CNN model, 924 ± 267,
indicating a substantial improvement in predictive accuracy.
Furthermore, the comparative analysis shows that DiffBatt consistently performs better than other
advanced models across various datasets. On the MIX dataset, DiffBatt achieved a mean RMSE of
202 ± 6, which is slightly higher than the best-performing model’s RMSE of 197 but significantly
lower than those of the deep learning models. Although DiffBatt did not achieve the lowest RMSE
on the MATR2 dataset (235 ± 16), it remains competitive compared to other deep learning models.
Importantly, DiffBatt exhibits a mean RMSE of 196 across all datasets, outperforming all other
models and demonstrating superior generalizability. These results illustrate DiffBatt’s efficacy in
learning and generalizing from diverse data sources. By utilizing the data splits and benchmarking
against results from BatteryML [47], we ensure that our comparisons are both fair and indicative of
DiffBatt’s capabilities.
In Fig. 3, we present the results for the RUL task for each test sample from the MIX dataset, obtained
using the DiffBatt model. Additionally, the figure includes SOH curves that correspond to the
predictions with the lowest and highest uncertainty. DiffBatt is capable of quantifying the uncertainty
in its predictions, which is reported here as the standard deviation of the RUL, computed from ten
generated samples. The data reveals that samples with higher prediction errors generally tend to
exhibit larger deviations in RUL.
In practical applications, estimating a battery’s SOH requires predicting the current discharge capacity
under standardized conditions using reference performance tests (RPTs) and historical cycling
data. However, the discrepancy between real-world battery usage and these standardized conditions
poses significant challenges in obtaining precise ground-truth labels, complicating accurate SOH
estimation. Consequently, developing a robust benchmark test for SOH estimation remains an
6
Table 1: Results obtained from DiffBatt for RUL prediction against benchmark results reported in
[47]. For models sensitive to initialization, values in the table correspond to the mean error and the
standard deviation, meanstd , across ten seeds. Values that are underlined represent the best results
from the benchmarks in BatteryML, while values that are bold indicate the overall best results.
Models MATR1 MATR2 HUST SNL CLO CRUH CRUSH MIX Mean
"Variance" model 136 211 398 360 179 118 506 521 304
"Discharge" model 329 149 322 267 143 76 >1000 >1000 411
"Full" model 167 >1000 335 433 138 93 >1000 331 437
Ridge regression 116 184 >1000 242 169 65 >1000 372 268
PCR 90 187 435 200 197 68 560 376 264
PLSR 104 181 431 242 176 60 535 383 264
Gaussian process 154 224 >1000 251 204 115 >1000 573 440
XGBoost 334 799 395 547 215 119 330 205 368
Random forest 1689 2337 3687 53225 1922 811 4165 1970 273
MLP 1493 27527 4599 37081 1465 1034 5659 45142 263
CNN 10294 228104 46575 924267 >1000 17492 54511 272101 464
LSTM 11911 21933 44329 53940 22212 10510 51939 2689 304
Transformer 13513 36425 39111 42423 18714 818 55021 27116 300
DiffBatt (ours) 884 23516 36823 12511 14014 1196 29418 2026 196
ongoing effort within the research community. Therefore, we benchmark our SOH estimation results
by approximating the SOH under actual operating conditions. Results are reported in Table 2. We
employ the same datasets in our RUL prediction tasks for the SOH estimation experiments, which
allows for reproducibility due to the clear data splits. For each degradation curve, we calculate
the error as the RMSE between the reference and predicted curves up to an SOH representing the
end of life (EOL). The reference SOH is zero-padded to match the length of the predicted curve
if necessary. Results are reported as the mean RMSE across test samples for different EOLs. To
ensure reproducibility, a detailed discussion on experimental setup and error metrics is included in
Appendix A.2.
σRUL
198 100
SOH(%)
2000
Prediction
148
99 90 Samples
1000 Reference
49 Prediction
0
80
0 0 500 1000 1500 0 500 1000 1500 2000
0 1000 2000
DiffBatt demonstrates strong performance in SOH estimation, achieving low RMSE values across
various datasets. When considering an EOL of 80%, our model exhibits the highest precision with
the CRUH dataset, yielding the lowest RMSE of 1.26 ± 0.04, and maintains consistent accuracy with
RMSE values of 1.68 ± 0.08 and 1.89 ± 0.03 on the MATR1 and SNL datasets, respectively. DiffBatt
has a higher RMSE with the HUST dataset at 2.75 ± 0.15, and results in an RMSE of 2.17 ± 0.07
with the MATR2 dataset. The MIX dataset, aggregating diverse data sources, results in an RMSE of
1.98 ± 0.03, indicating the model’s ability to generalize effectively. Consistent RMSE values across
the CLO (2.32 ± 0.07) and CRUSH (2.35 ± 0.05) datasets further highlight DiffBatt’s robustness across
different battery chemistries and operating conditions. To account for the varying EOL requirements
7
across different battery applications, such as electric vehicles and storage systems, we also evaluated
DiffBatt’s performance for EOL values of 90%, 70%, and 60%. At an EOL of 90%, DiffBatt
maintains strong precision with low RMSE values across all datasets. Despite the increased difficulty,
DiffBatt demonstrates robust performance with a minor increase in RMSE values across different
EOL thresholds, resulting in mean RMSE values of 1.19, 2.05, 2.59, and 3.16 for EOLs of 90%,
80%, 70%, and 60%, respectively. These results underscore DiffBatt’s adaptability for reliable SOH
estimations tailored to specific battery application requirements.
Table 2: Results obtained from DiffBatt for SOH estimation task considering different end of life
(EOL) values. Reported results correspond to the mean RMSE and the standard deviation, meanstd ,
across ten seeds.
EOL MATR1 MATR2 HUST SNL CLO CRUH CRUSH MIX Mean
90% 1.00 0.07 1.36 0.07 1.46 0.1 0.98 0.05 1.47 0.06 0.68 0.02 1.39 0.03 1.21 0.03 1.19
80% 1.68 0.08 2.17 0.07 2.75 0.15 1.89 0.03 2.32 0.07 1.26 0.04 2.35 0.05 1.98 0.03 2.05
70% 1.86 0.08 2.83 0.07 3.13 0.16 2.40 0.03 2.77 0.07 2.00 0.06 3.25 0.09 2.51 0.03 2.59
60% 2.23 0.08 3.97 0.07 3.43 0.16 2.95 0.04 3.16 0.09 2.43 0.07 3.97 0.12 3.13 0.02 3.16
8
Table 3: Results for the MATR1 dataset obtained using random
forests trained on synthetic data generated with different diffu-
sion guidance w, against results reported in [47] obtained from
a random forest trained without synthetic data.
w
Models RF [47] 0.0 1.0 2.0 4.0 6.0
FID (↓) NA 0.405 0.408 0.409 0.411 0.413
Precision (↑) NA 0.998 0.99 0.985 0.963 0.968
Recall (↑) NA 0.663 0.663 0.614 0.446 0.398
RMSE 1689 1096 1076 1046 1057 1068
Expressivity. DiffBatt is trained on the largest publicly available battery degradation datasets,
encompassing a wide range of battery chemistries and operational conditions. This comprehensive
training equips DiffBatt with a robust understanding of battery degradation patterns, making it an
ideal candidate for foundational modeling. Our results demonstrate DiffBatt’s high expressivity and
its ability to capture the complex dynamics involved in battery degradation.
Multimodality. In DiffBatt, the input conditions are encoded using a transformer encoder that can
be fine-tuned with relatively small amounts of new data from different cell chemistries. This ensures
the model’s quick adaptation to new battery technologies with minimal retraining. Furthermore,
additional data modalities such as temperature, current profiles, and environmental conditions can
be encoded and added to the condition vector, enhancing multimodality. This multimodal approach
ensures all relevant factors influencing battery health are considered, improving the model’s accuracy
and adaptability.
Scalability and memory. DiffBatt’s flexible architecture, leveraging diffusion models and trans-
formers, ensures efficient scalability and memory usage. Its design allows it to handle extensive
volumes of data and complex battery degradation scenarios effectively, making it capable of scaling
with increasing data inputs without compromising performance. Additionally, DiffBatt’s ability to
synthesize battery degradation curves furthers its versatility by enabling the creation of large, diverse
datasets. By generating synthetic degradation data, DiffBatt can augment existing datasets from cell
chemistries with limited samples, enhancing the generalization capabilities of downstream models.
9
5 Conclusions and outlook
Tackling battery degradation is a major hurdle in advancing green technologies and sustainable energy
solutions. Accurately predicting battery capacity loss remains particularly challenging due to its
intricate and complex nature. To address this issue, we present DiffBatt, a novel general-purpose
model for predicting and synthesizing battery degradation patterns based on diffusion models with
classifier-free guidance and transformer encoders. A key innovation is the integration of conditional
and unconditional diffusion models, enabling the robust generation of high-quality degradation
curves. DiffBatt functions as both a probabilistic model to capture the inherent uncertainties in aging
processes and a generative model to simulate and predict battery degradation over time.
We evaluate the performance of DiffBatt across three different tasks, i.e., RUL prediction, SOH
estimation, and SOH synthesis. In the RUL prediction task, DiffBatt achieved the lowest RMSE on
the MATR1, SNL, and CRUSH datasets, with RMSE values of 88±4, 125±11, and 294±18 respectively.
Notably, DiffBatt results in a mean RMSE of 196 across all datasets, significantly outperforming all
other competing models. These results illustrate DiffBatt’s efficacy in learning from and generalizing
across diverse data sources. Consistently low RMSE values in the SOH prediction task further
highlight DiffBatt’s robustness and reliability across different battery chemistries. Moreover, we
showcase the broad applicability of DiffBatt to generate high-quality battery degradation curves. We
show that augmenting battery datasets with synthetic data can lead to a better and more accurate
performance of downstream ML models, e.g., for RUL prediction.
We believe that by training on several diverse battery datasets and demonstrating strong generalizabil-
ity and robustness across various tasks, DiffBatt offers a promising pathway toward developing a
foundational model for battery degradation. However, to support a deep understanding of degradation
mechanisms and derive counter measures in battery design and or battery production as well as
formation the data-driven DiffBatt can be linked to physical-based models or needs to be extended
with regard to the variation of battery design and process parameters.
Acknowledgement
This research was conducted within the Research Training Group CircularLIB, supported by the
Ministry of Science and Culture of Lower Saxony with funds from the program [Link]
of the Volkswagen Foundation (MWK | ZN3678). We would like to thank the anonymous reviewers
who contributed to the improvement of our manuscript with their valuable suggestions and comments.
We acknowledge that Fig. 1 is adapted and modified from the work by Fürrutter et al. [14], available
on arXiv [15] under a CC BY-SA 4.0 license. The modifications include context-specific changes.
References
[1] P. M. Attia, K. A. Severson, and J. D. Witmer. Statistical learning for accurate and interpretable battery
lifetime prediction. Journal of The Electrochemical Society, 168(9):090547, 2021. doi: 10.1149/1945-7111/
ac2704. URL [Link]
[2] J.-H. Bastek, W. Sun, and D. M. Kochmann. Physics-informed diffusion models. arXiv preprint
arXiv:2403.14404, 2024. URL [Link]
[3] S. Blömeke, C. Scheller, F. Cerdas, C. Thies, R. Hachenberger, M. Gonter, C. Herrmann, and T. S.
Spengler. Material and energy flow analysis for environmental and economic impact assessment of
industrial recycling routes for lithium-ion traction batteries. Journal of Cleaner Production, 377:134344,
12 2022. ISSN 09596526. doi: 10.1016/[Link].2022.134344. URL [Link]
com/retrieve/pii/S0959652622039166.
[4] G. E. Blomgren. The development and future of lithium ion batteries. Journal of The Electrochemical
Society, 164(1):A5019–A5025, 12 2017. ISSN 0013-4651. doi: 10.1149/2.0251701jes. URL https:
//[Link]/article/10.1149/2.0251701jes.
[5] R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg,
A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel,
J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Durmus, S. Ermon, J. Etchemendy, K. Ethayarajh,
L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K. Goel, N. Goodman, S. Grossman, N. Guha, T. Hashimoto,
P. Henderson, J. Hewitt, D. E. Ho, J. Hong, K. Hsu, J. Huang, T. Icard, S. Jain, D. Jurafsky, P. Kalluri,
S. Karamcheti, G. Keeling, F. Khani, O. Khattab, P. W. Koh, M. Krass, R. Krishna, R. Kuditipudi, A. Kumar,
10
F. Ladhak, M. Lee, T. Lee, J. Leskovec, I. Levent, X. L. Li, X. Li, T. Ma, A. Malik, C. D. Manning,
S. Mirchandani, E. Mitchell, Z. Munyikwa, S. Nair, A. Narayan, D. Narayanan, B. Newman, A. Nie, J. C.
Niebles, H. Nilforoshan, J. Nyarko, G. Ogut, L. Orr, I. Papadimitriou, J. S. Park, C. Piech, E. Portelance,
C. Potts, A. Raghunathan, R. Reich, H. Ren, F. Rong, Y. Roohani, C. Ruiz, J. Ryan, C. Ré, D. Sadigh,
S. Sagawa, K. Santhanam, A. Shih, K. Srinivasan, A. Tamkin, R. Taori, A. W. Thomas, F. Tramèr, R. E.
Wang, W. Wang, B. Wu, J. Wu, Y. Wu, S. M. Xie, M. Yasunaga, J. You, M. Zaharia, M. Zhang, T. Zhang,
X. Zhang, Y. Zhang, L. Zheng, K. Zhou, and P. Liang. On the opportunities and risks of foundation models.
arXiv preprint arXiv:2108.07258, 2022. URL [Link]
[6] A. Brock, J. Donahue, and K. Simonyan. Large scale GAN training for high fidelity natural image synthesis.
arXiv preprint arXiv:1809.11096, 2019. URL [Link]
[7] M. Broussely, P. Biensan, F. Bonhomme, P. Blanchard, S. Herreyre, K. Nechev, and R. Staniewicz. Main
aging mechanisms in Li-ion batteries. Journal of Power Sources, 146(1-2):90–96, 8 2005. ISSN 03787753.
doi: 10.1016/[Link].2005.03.172. URL [Link]
S0378775305005082.
[8] P. V. Chombo and Y. Laoonual. A review of safety strategies of a Li-ion battery. Journal of Power
Sources, 478:228649, 12 2020. ISSN 03787753. doi: 10.1016/[Link].2020.228649. URL https:
//[Link]/retrieve/pii/S0378775320309538.
[9] F.-A. Croitoru, V. Hondru, R. T. Ionescu, and M. Shah. Diffusion models in vision: A survey. IEEE
Transactions on Pattern Analysis and Machine Intelligence, 45(9):10850–10869, 2023. doi: 10.1109/
TPAMI.2023.3261988. URL [Link]
[10] Z. Cui, L. Wang, Q. Li, and K. Wang. A comprehensive review on the state of charge estimation for
lithium-ion battery based on neural network. International Journal of Energy Research, 46(5):5423–5440,
4 2022. ISSN 0363-907X. doi: 10.1002/er.7545. URL [Link]
1002/er.7545.
[11] A. Devie, G. Baure, and M. Dubarry. Intrinsic variability in the degradation of a batch of commercial
18650 lithium-ion cells. Energies, 11(5), 2018. ISSN 1996-1073. doi: 10.3390/en11051031. URL
[Link]
[12] P. Dhariwal and A. Q. Nichol. Diffusion models beat GANs on image synthesis. In Advances in Neural
Information Processing Systems, 2021. URL [Link]
[13] J. S. Edge, S. O’Kane, R. Prosser, N. D. Kirkaldy, A. N. Patel, A. Hales, A. Ghosh, W. Ai, J. Chen, J. Yang,
S. Li, M.-C. Pang, L. Bravo Diaz, A. Tomaszewska, M. W. Marzook, K. N. Radhakrishnan, H. Wang,
Y. Patel, B. Wu, and G. J. Offer. Lithium ion battery degradation: what you need to know. Physical
Chemistry Chemical Physics, 23(14):8200–8221, 4 2021. ISSN 1463-9076. doi: 10.1039/D1CP00359C.
URL [Link]
[14] F. Fürrutter, G. Muñoz-Gil, and H. J. Briegel. Quantum circuit synthesis with diffusion models. Nature
Machine Intelligence, 6(5):515–524, May 2024. ISSN 2522-5839. doi: 10.1038/s42256-024-00831-9.
URL [Link]
[15] F. Fürrutter, G. Muñoz-Gil, and H. J. Briegel. Quantum circuit synthesis with diffusion models. arXiv
preprint arXiv:2311.02041, 2024. URL [Link]
[16] M. Galeotti, L. Cinà, C. Giammanco, S. Cordiner, and A. Di Carlo. Performance analysis and SOH (state of
health) evaluation of lithium polymer batteries through electrochemical impedance spectroscopy. Energy,
89:678–686, 9 2015. ISSN 03605442. doi: 10.1016/[Link].2015.05.148. URL [Link]
[Link]/retrieve/pii/S0360544215007756.
[17] R. Ginster, S. Blömeke, J. Popien, C. Scheller, F. Cerdas, C. Herrmann, and T. S. Spengler. Circular battery
production in the EU: Insights from integrating life cycle assessment into system dynamics modeling
on recycled content and environmental impacts. Journal of Industrial Ecology, 28(5):1165–1182, 10
2024. ISSN 1088-1980. doi: 10.1111/jiec.13527. URL [Link]
10.1111/jiec.13527.
[18] W. He, N. Williard, M. Osterman, and M. Pecht. Prognostics of lithium-ion batteries based on Dempster–
Shafer theory and the Bayesian Monte Carlo method. Journal of Power Sources, 196(23):10314–10321,
2011. doi: [Link] URL [Link]
science/article/pii/S0378775311015400.
[19] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. GANs trained by a two time-
scale update rule converge to a local Nash equilibrium. In Advances in Neural Information Processing
Systems, volume 30, 2017. URL [Link]
file/[Link].
11
[20] J. Ho and T. Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022. URL
[Link]
[21] J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information
Processing Systems, volume 33, 2020. URL [Link]
paper/2020/file/[Link].
[22] J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans. Cascaded diffusion models
for high fidelity image generation. Journal of Machine Learning Research, 23(47):1–33, 2022. URL
[Link]
[23] J. Hong, D. Lee, E.-R. Jeong, and Y. Yi. Towards the swift prediction of the remaining useful life of
lithium-ion batteries with end-to-end deep learning. Applied Energy, 278:115646, 2020. ISSN 0306-
2619. doi: [Link] URL [Link]
science/article/pii/S0306261920311429.
[24] D. Juarez-Robles, J. A. Jeevarajan, and P. P. Mukherjee. Degradation-safety analytics in lithium-ion cells:
Part I. aging under charge/discharge cycling. Journal of The Electrochemical Society, 167(16):160510, nov
2020. doi: 10.1149/1945-7111/abc8c0. URL [Link]
[25] D. Juarez-Robles, S. Azam, J. A. Jeevarajan, and P. P. Mukherjee. Degradation-safety analytics in lithium-
ion cells and modules: Part III. aging and safety of pouch format cells. Journal of The Electrochemical
Society, 168(11):110501, nov 2021. doi: 10.1149/1945-7111/ac30af. URL [Link]
1149/1945-7111/ac30af.
[26] D. P. Kingma and P. Dhariwal. Glow: Generative flow with invertible 1x1 convolutions. In Advances in
Neural Information Processing Systems, volume 31, 2018. URL [Link]
paper_files/paper/2018/file/[Link].
[27] T. Kynkäänniemi, T. Karras, S. Laine, J. Lehtinen, and T. Aila. Improved precision and
recall metric for assessing generative models. In Advances in Neural Information Process-
ing Systems, volume 32, 2019. URL [Link]
[Link].
[28] A. G. Li, A. C. West, and M. Preindl. Towards unified machine learning characterization of lithium-
ion battery degradation across multiple levels: A critical review. Applied Energy, 316:119030, 6 2022.
ISSN 03062619. doi: 10.1016/[Link].2022.119030. URL [Link]
retrieve/pii/S0306261922004354.
[29] T. Li, L. Biferale, F. Bonaccorso, M. A. Scarpolini, and M. Buzzicotti. Synthetic Lagrangian turbulence by
generative diffusion models. Nature Machine Intelligence, 6(4):393–403, Apr 2024. ISSN 2522-5839. doi:
10.1038/s42256-024-00810-0. URL [Link]
[30] W. Li, N. Sengupta, P. A. Dechent, D. Howey, A. Annaswamy, and D. U. Sauer. One-shot battery
degradation trajectory prediction with deep learning. Journal of Power Sources, page 230024, 2021. ISSN
0378-7753. doi: 10.1016/[Link].2021.230024. URL [Link]
record/820366.
[31] X. Li, D. Yu, V. Søren Byg, and S. Daniel Ioan. The development of machine learning-based remain-
ing useful life prediction for lithium-ion batteries. Journal of Energy Chemistry, 82:103–121, 7 2023.
ISSN 20954956. doi: 10.1016/[Link].2023.03.026. URL [Link]
retrieve/pii/S2095495623001870.
[32] C. Luo, Z. Zhang, S. Zhu, and Y. Li. State-of-health prediction of lithium-ion batteries based on diffusion
model with transfer learning. Energies, 16(9):3815, 4 2023. ISSN 1996-1073. doi: 10.3390/en16093815.
URL [Link]
[33] L. Luzi, P. M. Mayer, J. Casco-Rodriguez, A. Siahkoohi, and R. G. Baraniuk. Boomerang: Local
sampling on image manifolds using diffusion models. arXiv preprint arXiv:2210.12100, 2024. URL
[Link]
[34] G. Ma, S. Xu, B. Jiang, C. Cheng, X. Yang, Y. Shen, T. Yang, Y. Huang, H. Ding, and Y. Yuan. Real-
time personalized health status prediction of lithium-ion batteries using deep transfer learning. Energy
& Environmental Science, 2022. doi: 10.1039/D2EE01676A. URL [Link]
D2EE01676A.
[35] S. E. J. O’Kane, W. Ai, G. Madabattula, D. Alonso-Alvarez, R. Timms, V. Sulzer, J. S. Edge, B. Wu,
G. J. Offer, and M. Marinescu. Lithium-ion battery degradation: how to model it. Physical Chemistry
Chemical Physics, 24(13):7909–7922, 3 2022. ISSN 1463-9076. doi: 10.1039/D2CP00417H. URL
[Link]
[36] Y. Preger, H. M. Barkholtz, A. Fresquez, D. L. Campbell, B. W. Juba, J. Romàn-Kustas, S. R. Ferreira,
and B. Chalamala. Degradation of commercial lithium-ion cells as a function of chemistry and cycling
12
conditions. Journal of The Electrochemical Society, 167(12):120532, sep 2020. doi: 10.1149/1945-7111/
abae37. URL [Link]
[37] H. Rauf, M. Khalid, and N. Arshad. Machine learning in state of health and remaining useful life
estimation: Theoretical and technological development in battery degradation modelling. Renewable and
Sustainable Energy Reviews, 156:111903, 3 2022. ISSN 13640321. doi: 10.1016/[Link].2021.111903. URL
[Link]
[38] Z. Ren and C. Du. A review of machine learning state-of-charge and state-of-health estimation algorithms
for lithium-ion batteries. Energy Reports, 9:2993–3021, 12 2023. ISSN 23524847. doi: 10.1016/[Link].
2023.01.108. URL [Link]
[39] H. Rubenbauer and S. Henninger. Definitions and reference values for battery systems in electrical power
grids. Journal of Energy Storage, 12:87–107, 8 2017. ISSN 2352152X. doi: 10.1016/[Link].2017.04.004.
URL [Link]
[40] K. A. Severson, P. M. Attia, N. Jin, N. Perkins, B. Jiang, Z. Yang, M. H. Chen, M. Aykol, P. K. Herring,
D. Fraggedakis, M. Z. Bazant, S. J. Harris, W. C. Chueh, and R. D. Braatz. Data-driven prediction of
battery cycle life before capacity degradation. Nature Energy, 4(5):383–391, 3 2019. ISSN 2058-7546. doi:
10.1038/s41560-019-0356-8. URL [Link]
[41] J. Vetter, P. Novák, M. Wagner, C. Veit, K.-C. Möller, J. Besenhard, M. Winter, M. Wohlfahrt-Mehrens,
C. Vogler, and A. Hammouche. Ageing mechanisms in lithium-ion batteries. Journal of Power Sources,
147(1-2):269–281, 9 2005. ISSN 03787753. doi: 10.1016/[Link].2005.01.006. URL https://
[Link]/retrieve/pii/S0378775305000832.
[42] V. I. voor Technologisch Onderzoek (VITO), C. à l’Energie Atomique et aux Energies Al-
ternatives (CEA), Siemens, T. U. M. (TUM), T. S. B. Testing, ALGOLiON, R. A. Uni-
versity, L. Smart, T. U. Eindhoven, Voltia, and V. E. T. Solutions. Everlasting: Elec-
tric vehicle enhanced range, lifetime and safety through ingenious battery management,
2021. URL [Link]
Range_Lifetime_And_Safety_Through_INGenious_battery_management_/5065445/11.
[43] S. Wang, S. Jin, D. Bai, Y. Fan, H. Shi, and C. Fernandez. A critical review of im-
proved deep learning methods for the remaining useful life prediction of lithium-ion batter-
ies. Energy Reports, 7:5562–5574, 2021. ISSN 23524847. doi: 10.1016/[Link].2021.08.
182. URL files/1917/Acriticalreviewofimproveddeeplearningmethodsfortheremaining.
pdf[Link]
[44] Y. Xing, E. W. Ma, K.-L. Tsui, and M. Pecht. An ensemble model for predicting the remaining useful
performance of lithium-ion batteries. Microelectronics Reliability, 53(6):811–820, 2013. doi: [Link]
org/10.1016/[Link].2012.12.003. URL [Link]
pii/S0026271412005227.
[45] Y. Yang, M. Jin, H. Wen, C. Zhang, Y. Liang, L. Ma, Y. Wang, C. Liu, B. Yang, Z. Xu, J. Bian, S. Pan,
and Q. Wen. A survey on diffusion models for time series and spatio-temporal data. arXiv preprint
arXiv:2404.18886, 2024. URL [Link]
[46] B. Zhang, P. Xu, X. Chen, and Q. Zhuang. Generative quantum machine learning via denoising diffusion
probabilistic models. Physical Review Letters, 132:100602, Mar 2024. doi: 10.1103/PhysRevLett.132.
100602. URL [Link]
[47] H. Zhang, X. Gui, S. Zheng, Z. Lu, Y. Li, and J. Bian. BatteryML: An open-source platform for machine
learning on battery degradation. In The Twelfth International Conference on Learning Representations,
2024. URL [Link]
13
A Appendix
A.1 Datasets
Battery degradation curves utilized for training and testing the models are depicted in Fig. 4 for each
cell chemistry. A brief summary of the datasets included in BatteryML [47] is presented here.
The CALCE dataset includes full lifecycle data from 13 batteries with an LCO cathode. Each battery
has a nominal capacity of 1100 mAh. They were all charged using a constant current/constant voltage
protocol: 0.5C current until reaching 4.2V, maintaining 4.2V until the current dropped below 0.05A,
and a cutoff voltage of 2.7V [44, 18].
The MATR dataset, provided by Severson et al. [40] and Hong et al. [23], is one of the largest
public datasets containing 180 commercial 18650 LFP batteries. These batteries, cycled at a forced
convection temperature chamber of 30◦ C, have a nominal capacity of 1.1 Ah and a nominal voltage
of 3.3V. The dataset comprises three subsets: MATR1, MATR2 [40], and CLO [23], all categorized due
to distinct measurement batches.
The HUST dataset includes 77 LFP batteries, similar to those in the MATR dataset. These batteries
followed an identical charging protocol with varying multi-stage discharge protocols, all conducted
at a constant temperature of 30◦ C [34].
The HNEI dataset contains 14 commercial 18650 cells with a graphite anode and a blended NMC and
LCO cathode. These cells were cycled at 1.5C to 100% depth of discharge for over 1000 cycles at
room temperature [11].
The SNL dataset includes 61 commercial 18650 cells (NCA, NMC, and LFP), cycled to 80% capacity.
The study evaluates the impact of temperature, depth of discharge, and discharge current on long-term
degradation [36].
The UL_PUR dataset comprises 10 commercial pouch cells with a graphite negative electrode and an
NCA cathode. These cells were cycled at 1C between 2.7V and 4.2V, equivalent to 0-100% state of
charge (SOC), at room temperature until reaching 10-20% capacity fade. Additionally, modules were
cycled at C/2 between 13.7V and 21.0V until 20% capacity fade [24, 25].
The RWTH dataset contains data from 48 lithium-ion battery cells aged under identical conditions.
These cells feature a carbon anode and an NMC cathode [30]. The cells were cycled at a constant
ambient temperature of 25◦ C. Each cycle involved a 30-minute discharge phase down to 3.5V and a
30-minute charge phase up to 3.9V, with the currents capped at a maximum of 4A. This resulted in
cycles between approximately 20% and 80% state of charge.
100
80
60
100
80
60
600 1200 1000 2000 1000 2000 500 1000 400 800
Figure 4: Train (up) and test (bottom) samples for each cell chemistry. The data is scaled using the
SOH of the first cycle.
For prediction tasks, i.e. we employ a guidance strength of w = 0.0 and generate ten samples for
each input capacity matrix. The capacity matrix is constructed from the first 100 cycles. We further
select the final prediction based on the best fit to the SOH of the first 100 cycles. The RMSE for an
SOH sample j is computed as
v
u nj
u1 X
RMSEj = t (ỹi − yi )2 (8)
nj i=1
14
where ỹ and y represent the predicted and the reference SOH in percentage, respectively, i denotes
the cycle number and nj is the cycle number at which the predicted SOH reaches the EOL. Further,
we report the mean RMSE across all the test samples as the RMSE for the dataset.
Figure 5 illustrates the predicted SOH versus the reference SOH for all test samples in the MIX
dataset. The results demonstrate that DiffBatt effectively captures various degradation dynamics
and accurately predicts SOH for the majority of test samples and highlights DiffBatt’s ability to
generalize across different battery chemistries and operational conditions present in the MIX dataset.
This capability is essential for developing reliable battery health monitoring systems that can adapt to
diverse usage patterns and environmental factors.
SOH(%)
100
80
100
80
100
80
100
80
100
80
100
80
100
80
100
80
100
80
100
80
100
80
100
80
Cycle
Figure 5: SOH predictions against reference for all the test samples of MIX dataset. The pink dashed
line shows the prediction and the cyan solid line shows the reference.
100
90
0 500 1000 1500 2000 0 500 1000 1500 2000 0 500 1000 1500 2000 0 500 1000 1500 2000 0 500 1000 1500 2000
Figure 6: MATR1 synthetic data generated by DiffBatt with different guidance strengths.
15
DiffBatt quantifies prediction uncertainty by calculating the standard deviation of the RUL from ten generated samples, revealing that samples with higher prediction errors tend to show larger deviations in RUL. This capability is evidenced by the range of RUL uncertainties for different test samples from the MIX dataset .
DiffBatt demonstrates superior performance in RUL prediction tasks by achieving the lowest RMSE on the MATR1, SNL, and CRUSH datasets with RMSE values of 88 ± 4, 125 ± 11, and 294 ± 18 respectively. On the MIX dataset, DiffBatt's RMSE was slightly higher at 202 ± 6 compared to the best-performing model at 197, but it significantly outperformed deep learning models. Overall, DiffBatt's mean RMSE across all datasets was 196, outperforming all other models .
Practical challenges in estimating a battery's SOH include the discrepancy between standardized conditions of reference performance tests and real-world battery usage, which complicates obtaining precise ground-truth labels for accurate SOH estimation .
DiffBatt outperforms other models in SOH estimation by achieving low RMSE values across various datasets, such as 1.26 ± 0.04 for the CRUH dataset and 1.98 ± 0.03 on the MIX dataset. These results demonstrate DiffBatt's high precision and generalizability across different battery chemistries and conditions, indicating its strong capability in accurate SOH predictions .
DiffBatt exhibits varying performance when evaluated for different EOL percentages. At an EOL of 90%, RMSE remains low, indicating high precision across datasets. As the EOL decreases to 80%, 70%, and 60%, there are marginal increases in RMSE, but it still maintains robust performance. This adaptability implies that DiffBatt can provide reliable SOH predictions tailored to specific application requirements, such as in electric vehicles or energy storage systems .
Data augmentation in the DiffBatt model is implemented by generating synthetic SOH curves, which enhances limited battery degradation datasets. This process helps improve the model's training by providing a richer dataset for machine learning tasks, thus enabling better performance in predicting battery degradation .
DiffBatt ensures fair comparisons and consistent performance by utilizing data splits provided by BatteryML, benchmarking against their results, and employing standardized evaluation metrics like RMSE across various datasets. This approach allows for a fair, reproducible comparison of model performance on standardized benchmarks .
DiffBatt adapts its performance metrics by maintaining strong precision with low RMSE values across different EOL thresholds. For EOL values of 90%, 80%, 70%, and 60%, DiffBatt's mean RMSE values were 1.19, 2.05, 2.59, and 3.16, respectively, illustrating its robustness for SOH estimation under varying application-specific EOL conditions .
SOH synthesis involves generating new state of health curves for data augmentation purposes. DiffBatt uses this process to enrich limited battery degradation datasets for machine learning tasks, thereby enabling more robust model training and better predictive performance .
DiffBatt offers the advantage of lower RMSE in predicting RUL on various datasets, showcasing robustness and precision over other models. It consistently outperforms benchmarks like PCR and CNN, with significant improvements in predictive accuracy, flexibility across different battery compositions, and capability to generalize across diverse operational conditions .