Machine Learning for Downtime Analysis
Machine Learning for Downtime Analysis
*Correspondence:
kosta@[Link] Abstract
1
Faculty of Engineering, Manufacturing companies focus on improving productivity, reducing costs, and align-
University of Kragujevac, ing performance metrics with strategic objectives. In industries like paper manufactur-
Kragujevac, Serbia ing, minimizing equipment downtime is essential for maintaining high throughput.
2
Faculty of Natural Sciences
and Mathematics, University Leveraging the extensive data generated by these facilities offers opportunities
of Montenegro, Podgorica, for gaining competitive advantages through data-driven insights, revealing trends,
Montenegro patterns, and predicting future performance indicators like unplanned downtime
3
Faculty of Electrical
Engineering, University length, which is essential in optimizing maintenance and minimizing potential losses.
of Montenegro, Podgorica, This paper explores statistical and machine learning techniques for modeling down-
Montenegro time length probability distributions and correlation with machine vibration measure-
ments. We proposed a novel framework, employing advanced data-driven techniques
like artificial neural networks (ANNs) to estimate parameters of probability distributions
governing downtime lengths. Our approach specifically focuses on modeling param-
eters of these distribution, rather than directly modeling probability density function
(PDF) values, as is common in other approaches. Experimental results indicate a sig-
nificant performance boost, with the proposed method achieving up to 30% superior
performance in modeling the distribution of downtime lengths compared to alter-
native methods. Moreover, this method facilitates unsupervised training, making it
suitable for big data repositories of unlabelled data. The framework allows for potential
expansion by incorporating additional input variables. In this study, machine vibration
velocity measurements are selected for further investigation. The study underscores
the potential of advanced data-driven techniques to enables companies to make
better-informed decisions regarding their current maintenance practices and to direct
improvement programs in industrial settings.
Keywords: Lean industrial systems, Paper manufacturing, Production downtime, Big
data analytics, Machine learning, Unsupervised learning, Artificial neural networks,
Probability distribution, Parameter estimation, Maximum likelihood estimation
© The Author(s) 2024. Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0
International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long
as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you
modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of
it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise
in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted
by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy
of this licence, visit [Link]
Introduction
In today’s competitive global business environment, manufacturing companies face
increasing challenges in terms of productivity improvement, implementation of inno-
vative technologies, and more demanding environmental regulations.
To increase competitiveness performance-oriented companies, especially organiza-
tions operating in accordance with Lean principles [1] are focused on reducing the
most common causes of productivity losses called the six big losses [2] and establish-
ing clear performance indicators aligned with the strategic objectives throughout the
organization. Based on the business strategy, measurable success goals are defined to
direct the activities of the organization as well as measure achieved business results.
Over the past several decades, maintenance has become a critical factor for achiev-
ing business goals and maintaining competitiveness. To maximize productivity com-
panies must achieve the right performance from their production equipment which is
directly influenced by performed maintenance activities. Industrial case studies dem-
onstrate how applying effective maintenance practices can improve the productivity
and profitability of the production process by preventing production downtime and
inadequate product quality due to maintenance-related causes [3]. Therefore, mainte-
nance should not be seen as a cost center, but as a function that generates profit.
The directions of development of the maintenance function in the future changed
conditions of digitized production are not easy to predict. Changes such as the wider
application of advanced methods for data analysis, increased focus on education and
training of employees, new approaches to maintenance planning, and more demand-
ing environmental regulations have a crucial impact on development directions and
the future role of maintenance [4].
The development of industrial companies’ maintenance function depends to a large
extend on the growing trend of automation and application of advanced technologies
in production processes denoted by the term Industry 4.0. Industry 4.0 is transform-
ing industrial companies by embracing digitalization, automation, artificial intel-
ligence, big data, machine learning, cloud computing, and Internet of Things (IoT)
aiming for systems interacting with each other, autonomous decisions, and smart
factories.
With the exponential growth of data in recent years, industries have access to mas-
sive amounts of valuable information. Big data analytics helps to make sense of this
data and transform it into actionable insights for planning maintenance activities.
By integrating this two activities, companies can identify potential issues before they
occur, and ultimately enhance their overall performance.
The process of leveraging big data analytics for applying effective maintenance prac-
tices typically involves the following steps:
1. Data collection—Sensors and IoT devices are installed on machinery to collect vari-
ous types of data, including vibration, temperature, and oil analysis.
2. Data analysis—Interpretation of this big data using a combination of analysis tools
like automated machine learning and expert systems.
3. Alerts and notifications—Insights presented in dashboards or reports on the possible
consequences of the diagnosis and when the failure may occur.
The approach presented in this paper should enable key performance indicators (KPI)
prediction that give a picture of the equipment maintenance, based on the historical per-
formance levels. The dataset used for exploration has the following information with a
daily granularity: planned or scheduled downtime, unplanned mechanical equipment
downtime and unplanned electrical equipment downtime, together with paper machine
vibration measurements.
Therefore, this paper aims to explore the effectiveness of statistical and machine
learning methods for modeling probability distribution parameters in downtime length
analysis and joint probability distribution of recorded production downtime data and
introduced parameter, namely vibration measurements collected from permanently
mounted accelerometers on the paper machine. Based on the analysis of this informa-
tion, it is possible to better understand currently applied maintenance and business
practices and define the basic directions for more successful implementation of improve-
ment programs, including improvements based on performance-based maintenance.
The structure of this paper is based on the guidelines provided in [5, 6]. The second
section defines the problem of modeling unplanned machine downtime in industrial set-
ting and its implications regarding assessing most important performance KPIs. The role
of big data analytics in improving maintenance practices is highlighted and its signifi-
cance in minimizing downtime and maximizing productivity. In the third section previ-
ous approaches in modeling unplanned production downtime lengths are analyzed, and
the best current solution is identified. The fourth section presents fundamental com-
ponents of our proposed solution, its architectural design, and the research and meth-
odological approaches employed in its development. Following, fifth section, elaborates
on the proposed approach in detail, explaining statistical and machine learning methods
employed for modeling downtime lengths distributions, describing the dataset used for
the experiments, and providing comparison with the best-performing approach from
the literature. Final, concluding chapter summarizes the accomplishments and outlines
emerging research directions.
Problem definition
An emerging trend in manufacturing involves shifting focus from the mere cost of main-
tenance activities to emphasizing the value they deliver. Performance-based mainte-
nance is an approach to industrial systems maintenance where the achieved KPI results
are associated with structured incentives, rather than a process for achieving outcomes
[7]. Hence, forecasting future KPI results becomes crucial as it enables companies to
make better-informed decisions regarding their current maintenance practices and the
realistic achievements of performance-based maintenance improvement programs.
The pulp and paper industry are a growth market due to an increase in fiber-based
packaging as an alternative solution to plastic packaging. A paper manufacturing
machine, or paper machine, is used to manufacture paper in large quantities at high
speed. Modern paper machines can be more than 10 m wide, 20 m high, 400 m long,
and incorporate as many as 1500 bearings [8]. Paper machines differ in design according
Fig. 1 An overview of a typical papermaking process using recycled paper (1—pulper, 2—screen, 3—
flotation cells, 4—screen, 5—cleaner, 6—thickener, 7—refiner, 8—cleaner, 9—screen, 10—blend chest, 11—
fan pump, A—forming section, B—press section, C—dryer section, D—calender, E—reeler, F—winder)
to the grade of paper they are producing. Generally, they consist of a wire or forming
section, a press section, a drying section, a coating section, a calendar and a reeler, as
illustrated in Fig. 1.
Deinking and stock preparation processes, depicted in Fig. 1, ensure that recycled
paper meets defined quality criteria for fiber characteristics, additives, and contaminants
before being fed into the forming section of the paper machine, where the stock, initially
containing around 99% water, is processed to reduce its water content to about 80%.
Subsequently, in the press section, the paper web undergoes multiple press nips to fur-
ther expel water, resulting in a water content typically ranging between 50 and 65%. Dry-
ing in the dryer section then reduces the moisture content to 5%–10%. For coated paper,
additional calendering or glazing processes achieve a smoother, glossy finish. After pro-
duction, the large continuous paper web, typically 8–10 m wide, is moved to winder to
be cut into smaller rolls for distribution, with the initial reeling performed at the end of
the paper machine in the reeler.
Condition monitoring is essential to the paper manufacturing process to prevent cata-
strophic failures and ensure operational uptime. The aim of this process is to determine
the condition of components that influence machine reliability. Paper machines have a
wide variety of defects and conditions related to problems in bearings, gears, low-speed
and variable-speed operation and other adverse factors due to harsh conditions. Vibra-
tion analysis is the most used technique in paper machine condition monitoring. The
advantage of condition monitoring based on vibration analysis is that it acts as an early
warning system, providing time for maintenance activities planning.
Vibration measurements used for exploration in this paper are based on measure-
ments of accelerometers permanently mounted in different manufacturing process
sections. Table 1 presents the distribution of vibration monitoring system sensors in
different sections. Every roll, gearbox and motor on a paper machine has a vibration
sensor permanently mounted on the predefined measurement positions. Accelerom-
eters used for paper machine vibration measurements are mounted using threaded
studs as close as possible to the rolling bearings to faithfully record mechanical vibra-
tions and transform them into electrical signals for condition monitoring and mainte-
nance. Due to the often hot, wet, and chemically challenging environment present in
paper machine vibration sensors installation needs to be robust and tailored to meet
environmental conditions. All sensors have waterproof connectors, and sensor cables
are run through protective tubes and cable trays to stainless steel junction boxes situ-
ated up to 50 ms from its sensors. The location has been chosen so that it is con-
venient to connect the optimal number of sensor cables. In addition, situating these
junction boxes away from the machine avoids exposing them to the machine’s harsh
environment.
With more than 12,000 vibration Fast Fourier Transform (FFT) spectra recorded
daily, the dataset constitutes a vast amount of data. Three vibration parameters are
calculated for each sensor: vibration velocity, vibration acceleration and Svenska Kul-
lager Fabriken (SKF) acceleration enveloping [9]. Frequency spectrum parameters are
determined based on the rotational speed and the frequency of interest at each meas-
urement position, with measurements being scheduled. However, in the event of an
alarm, more frequent measurements are recorded. Figure 2 presents the number of
vibration measurements recorded daily over a certain period.
In this paper, we consider unplanned downtime periods, which are further cat-
egorized into those related to mechanical and electrical failures. Analyzed data will
consist of downtime lengths recorded by an automated paper machine downtime
detection system and classified based on the main downtime reason. Developed
statistical/machine learning models are used to estimate key factors affecting the
operational availability of machines, which is one of the three core parameters of the
overall equipment effectiveness (OEE).
OEE is the most used efficiency measure in industrial companies. The term was first
introduced as a component of Total Productive Maintenance concept [10]. This metric
identifies and categorizes major losses or reasons for poor asset performance and pro-
vides the basis for determining improvement priorities.
The three core parameters that has an affect OEE are:
1. Availability—Ratio of the time when equipment is available for production and total
time,
2. Performance—Ratio of the production pace or speed and the theoretical maximum
speed,
3. Quality—Ratio of the volume of final product with approved quality and total pro-
duction volume.
MTBM
Ao = · 100 (1)
MTBM + MDT
where MTBM represents “Mean Time Between Maintenance” and MDT is “Mean Down
Time”.
Given that planned downtime is scheduled and known in advance, our focus lies pre-
dominantly on modeling the aspect of availability pertaining to unplanned downtime
attributed to machine failures. Data recorded by automated paper machine downtime
detection system is divided based on the main downtime reasons: mechanical failures
and electrical failures, and it is subsequently used for exploration and modelling.
Existing solution(s)
In recent years the role of big data and machine learning in improving maintenance
practices and downtime minimization has been a subject of numerous academic and
industrial researchers. Generally, published papers could be classified in two research
areas, namely machine fault detection and prognosis and downtime, availability, and
OEE prediction as the most used efficiency measures in industrial settings.
The application of machine learning in machine fault detection and prognosis in real
industrial manufacturing is the subject of numerous academic and industrial research.
A comprehensive and systematic literature reviews [12, 13] present an overview of the
challenges faced when using machine learning methods to detect mechanical faults and
Proposed solution
This paper introduces a modeling framework specifically designed for capturing the
probability distribution of machine downtime lengths resulting from electrical or
mechanical failures. This framework could ultimately be used in the estimation of the
production time losses caused by such failures. Schematic representation of the frame-
work is presented in Fig. 3.
The data processing pipeline starts with data acquisition from two primary sources.
The first acquisition component is an automated machine downtime detection system
equipped with stop detection sensors strategically placed around the machine. Upon
detecting a machine stop, the system records the duration of the downtime. Subse-
quently, the type of the downtime is identified based on the cause that triggered it. The
second data acquisition component involves a vibration monitoring system, compris-
ing of more than 500 sensors distributed across the paper machine. Through frequency
domain analysis, this system extracts key vibration parameters such as vibration velocity,
vibration acceleration, and SKF acceleration enveloping. Measurement results used for
exploration and modeling are spectrum based overall values.
The acquired data undergoes further analysis to model its probability distribution. At
the core of the framework lies a parameter estimation model tailored for determining
parameters of the assumed probability distribution. Preceding the parameter estimation
model, the proposed framework involves statistical analysis of the distribution of the
available data, to identify parameters requiring estimation.
Traditional statistical approaches to parameter estimation have served as the baseline
for estimating distribution parameters but have been enhanced by the introduction of
ANNs. ANNs yield superior results in modeling distribution functions and their param-
eters due to their ability to capture intrinsic features of the data that cannot be effectively
captured by standard statistical methods alone. These features may include specific
operational characteristics of a given company and patterns of behavior during mainte-
nance interventions, which lead to diverse patterns in downtime.
The proposed framework eliminates the need for discretization of downtime lengths.
Additionally, it enables the development of more precise models compared to general
machine learning approaches agnostic of probability distribution or traditional statis-
tical estimations. Moreover, it facilitates unsupervised training, making it suitable for
repositories of unlabeled big data.
The framework allows for potential expansion by incorporating additional input vari-
ables. Modeling process is enriched by estimating parameters of the joint probability
distribution, i.e. conditional distribution of the introduced parameter and downtime
length. These additional variables could include various machine attributes, such as
the average speed of the machine, the velocity or acceleration of vibrations induced by
machine operation, as well as categorical attributes like the cause, type, or location of the
failure, provided such data is available. In this study, machine vibration velocity meas-
urements are selected for further investigation. The relationship between downtime
length and vibrations has not been previously explored in the literature, to the best of
our knowledge.
By analyzing the joint distribution, patterns or correlations between measured vibra-
tion levels and downtime length can be identified. The joint probability distribution
model could serve as a tool to recognize and respond to vibration patterns associated
with increased downtime. If certain vibration levels consistently correlate with higher
downtimes, this insight can help to uncover underlying machinery issues and guide
maintenance strategies, leading to more targeted and proactive maintenance efforts.
This information can be used to allocate resources more effectively, such as ensuring
the availability of key spare parts or scheduling additional personnel during anticipated
downtime periods.
The simulation component, depicted in Fig. 3, serves as the engine of the framework.
It generates failures over the specified time period, based on the estimated distribution.
This simulation serves as the foundation for estimating production time losses. By simu-
lating the occurrence of failures and their associated lengths, our framework facilitates
the estimation of production downtime, which offers valuable insights into the potential
impacts of machine failures on overall production efficiency.
The design of this framework was carried out using several scientific methods elab-
orated in [6]. It predominantly involved elements of hybridization (H). Employing
well-established statistical and machine learning methods helped in creating an ANN
that more accurately models the distribution compared to general machine learning
approaches. Additionally, it includes elements of specialization (S), as it leverages well-
established statistical and machine learning techniques for knowledge extraction within
the industrial production processes domain.
Elaboration
This section expands upon the specifics of the framework outlined in previous section. It
offers comprehensive overview of methodologies employed, the experimental setup, and
the achieved results.
Dataset
The dataset used for conducting experiments contains data regarding mechanical and
electrical failures occurring on a paper production machine, along with machine vibra-
tion measurements. Data was collected over the period spanning from December 2021
to January 2024. Besides downtime lengths and failure types, thousands of vibration
measurements were conducted daily on the machine, and the average daily vibration
velocity was selected for further analysis. It is important to note that errors can occur
during measurements, often resulting from moisture ingress in the vibration sensor
connector due to harsh operating conditions, resulting in unrealistically high measure-
ment results. To mitigate the impact of these errors, any measured values exceeding 50
mm/s were excluded from the daily average calculation as a part of outlier detection and
removal process.
The Z-score method is used to identify outliers in downtime lengths data by measur-
ing how many standard deviations an individual data point is away from the mean of the
dataset.
Data points with a Z-score greater than a certain threshold (here set to 3) are consid-
ered outliers. More information about outlier detection techniques and strategies for
managing outliers can be found in [21].
Data visualizations within this section were generated on the outlier-free dataset.
Boxplots in Fig. 4 provide a concise overview of mechanical and electrical failures.
Notably, the dominant downtime length is zero, indicating that most days are devoid
of machine failures. These are not shown in figures for clarity. The median down-
time length is approximately 50 min for both types of failures, with similar minimum
and maximum lengths. The minimum length of a downtime caused by a mechanical
failure is 5 min, while for an electrical failure it is 10 min. The maximum downtime
lengths are 363 and 320 min for mechanical and electrical failures, respectively.
Histograms of downtime lengths for both types of failures are presented in Fig. 5.
Each failure type is depicted with two histograms containing 10 and 20 bins, respec-
tively. Significant differences are observed from these charts, highlighting the chal-
lenges associated with discretizing the data space into intervals. The influence of
the chosen finite number of intervals on model predictions in [17] is thus further
emphasized.
The most recent 20% of the data, corresponding to approximately the last 5 months
of 2023, is left out for model evaluation purposes.
and
1 − e−x if x ≥ 0
F (x) =
0 otherwise, (4)
respectively.
Machine vibration velocity variable Y is also analyzed and its distribution is deter-
mined based on the available dataset. By examining the histogram of Y alongside the
gamma distribution, we can form an initial hypothesis that requires further investiga-
tion. The PDF of the gamma distribution G (k, θ ) is given by:
θ k k−1 −θ y
g(y) = y e , for y > 0, (5)
Ŵ(k)
where k > 0 is the shape parameter, θ > 0 is the rate parameter, and Ŵ(k) denotes the
gamma function.
The CDF of the gamma distribution is given by:
1
G(y) = γ (k, θ y), (6)
Ŵ(k)
Table 2 Estimated unknown parameters with SE of the gamma and bootstrap values of
Kolmogorov Smirnov test statistic Dboot and pboot value
Vibration k̂ (se) θ̂ (se) Dboot pboot
n
θ k k−1 −θ yi
L2 (k, θ; y1 , y2 , . . . , yn ) = y e .
Ŵ(k) i
i=1
When estimating distribution parameters, it is common practice to work with the log-
likelihood function since it simplifies calculations and is numerically more stable. This
function is derived by taking the natural logarithm of the likelihood function:
n
n
ℓ2 (k, θ ; y1 , y2 , . . . , yn ) = (k − 1) ln yi − θ yi + n k ln θ − n ln Ŵ(k). (7)
i=1 i=1
Fig. 6 Histogram and PDF of the daily average vibration velocity measurements
where η = 3ρ (|η| ≤ 1), (f (x), g(y)) and (F (x), G(y)) are the marginal PDFs and CDFs
of X and Y defined in Eqs. (3–5) and (4–6), respectively. ρ is the correlation coefficient
between X and Y , obtained by:
E(X − µX )(Y − µY )
ρ= , (10)
σX σY
where (µX , σX ) and (µY , σY ) are the population mean and standard deviation of X and Y ,
respectively. These parameters are often replaced by sample mean and sample standard
deviation.
where x1 , x2 , .., xn are the observed samples and n is the number of observed samples.
The estimator ˆ is obtained as a solution of the maximization problem:
1
ˆ = max l1 (; x1 , . . . , xn ) = , (12)
xn
where functions ℓ2 and ℓ1 are log-likelihood functions from (7) and (11), respectively, F
and G are respective CDFs, and η = 3ρ is the correlation coefficient. This likelihood is
calculated for a set of n randomly sampled observations (xi , yi ), i = 1, 2, ..., n from the
FGM bivariate gamma distribution with unknown parameters , k and θ.
2Ŵ (k,θyi )
∂ n
n n
ηxi e−xi Ŵ(k) −1
ℓ3 = − xi + 2Ŵ (k,θ yi ) = 0, (14)
∂
η 1 − e−xi −1 +1
i=1 i=1 Ŵ(k)
n
∂
ℓ3 = ln yi + n ln θ − nψ(k)+
∂k
i=1
−xi
3,0 1, 1
n η 1 − e 2 G 2,3 θy i | + ln θy i Ŵ k, θyi − 2ψ(k)Ŵ k, θyi
0, 0, k
2Ŵ (k,θ yi ) = 0,
i=1 Ŵ(k) η 1 − e−xi Ŵ(k) −1 +1
(15)
n n k−1
2ηyi 1 − e−xi e−θ yi θ yi
∂ nk
ℓ3 = − yi − 2Ŵ (k,θyi )
= 0, (16)
∂θ θ Ŵ(k) η 1 − e−xi −1 +1
i=1 i=1 Ŵ(k)
m,n z a1 , . . . , ap is the Meijer G-function and ψ(z) is digamma function.
where Gp,q
b1 , . . . , bq
As the system above equation does not have explicit solutions, in order to obtain the
ML estimates, we maximize the log-likelihood function through numerical optimization
procedure.
ANN estimation
The baseline estimation of distribution parameters is improved by training an ANN
estimator, illustrated in Fig. 7. Architecture of the ANN was designed with the aim to
maintain a comparable level of model complexity as in [17, 20]. The ANN consists of 2
hidden layers with 256 neurons each. At the input layer, the only mandatory variable is
the downtime length, while the others are optional. These additional inputs may include
x − min(x)
x′ = , (17)
max(x) − min(x)
where x′ is the scaled value, x is the original value, and min(x) and max(x) are minimum
and maximum values of the variable x in the dataset, respectively.
The activation function used in the hidden layers of the network is the sigmoid, while
the activation of the output layer is softplus. Softplus is defined with the following
equation:
f (x) = 1 + ln 1 + ex .
(18)
We opted for this function in the output layer because all considered parameters in both
exponential and gamma distribution take positive values, which corresponds with the
characteristics of the softplus function.
In training ANNs, we retained the maximum likelihood estimation approach from
the baseline model, and therefore used functions defined in (11) and (13) as optimiza-
tion objectives. As the ANN training typically involves minimization of a loss function,
we transform the log-likelihood functions to their negative log-likelihood equivalents.
Employing negative log-likelihood as the loss function enables an unsupervised training
procedure for the ANN, as negative log-likelihood is calculated directly from the data
without the need for empirical PDF values or other ground-truth labels.
As previously noted, the dataset contains the prevalence of zero values regarding
downtime lengths. This could potentially bias the model towards predicting non-fail-
ure scenarios. However, modeling downtime lengths greater than zero is of paramount
importance as these instances represent actual failure events in the production process
and directly impact productivity and efficiency. By focusing on more precise estima-
tions of longer downtime lengths, we ensure that the model accurately captures the most
impactful events. This approach better aligns with the practical goal of minimizing pro-
duction disruptions and optimizing maintenance schedules in industrial settings.
To avoid overfitting and to some extent mitigate the influence of the prevalence of
zero values on model predictions, ANN weights are regularized with the L2-norm. The
regularization parameter is set to 0.001, to balance between preventing overfitting and
preserving model flexibility. All ANNs were trained using the Adam optimizer with a
learning rate set to 0.001. Number of epochs for modeling exponential distribution is
set to 10, while the networks estimating joint-probability parameters are trained for 20
epochs.
Experimental results
We evaluated our baseline statistical model, ANN model, and the ANN from [17] on a
holdout dataset. Comparing our approach with that of [17] is not straightforward, given
that network in [17] was trained on data in the discrete domain, while our model oper-
ates in the continuous domain. To enable comparison, the PDF curve is interpolated
from discrete domain model [17] by non-linear least squares method [27]. The mid-
points of each interval into which the range of downtime lengths is divided in [17], along
with the corresponding PDF values generated by the proposed network, were used as
interpolation points. It is expected that for these points the ANN will provide the closest
PDF estimations. Initial guess for the curve fitting procedure is the curve obtained with
the baseline statistical model, i.e. the distribution parameters derived from that model.
Table 3 presents the negative log-likelihood (NLL) values obtained for each of the
models on a holdout dataset. NLL is most commonly used evaluation metric for proba-
bilistic models. It provides a measure of models’ ability to accurately capture patterns
in previously unseen data by quantifying the likelihood of the observed data under the
model. A lower NLL indicates a better fit of the model to the data. In the terms of our
models, we measure how well they capture the distribution of downtime lengths in the
dataset. The obtained results demonstrate that our ANN model outperforms baseline
Table 3 Negative log-likelihood values on a holdout set. Bold values indicate the model with the
best performance
Probability distribution Negative
log-
likelihood
Exponential
Mechanical failures
ANN 389.69
Baseline 394.34
[17] 453.12
Electrical failures
ANN 595.48
Baseline 709.98
[17] 599.40
Joint
Mechanical failures
ANN 771.50
Baseline 790.59
[17] 1089.37
Electrical failures
ANN 985.77
Baseline 1105.90
[17] 1003.15
Fig. 8 PDFs of downtime lengths caused by mechanical and electrical machine failures
Fig. 9 Joint PDFs of downtime length and vibration velocity (mechanical failures—first row, electrical
failures—second row)
In Fig. 9, which depicts joint PDFs generated by each of the models, it is noticeable
that comparative models neglect a larger part of the range, with PDFs concentrated
around the maximum value and extremely small values elsewhere. Conversely, the
PDF approximated with the ANN is much more spread out across the entire range.
The ANN model assigns higher probability for non-zero failure lengths compared to
the other approaches, which tend to converge faster towards zero as the downtime
length increases, underestimating the probability of failure lengths as they deviate
from zero.
Conclusion
This paper presents a comprehensive modeling framework for machine downtime
lengths resulting from mechanical and electrical failures in industrial settings. The
proposed method shows that modeling the distribution parameters offers significant
advantages in machine downtime length analysis over directly modeling PDF values.
Leveraging advanced techniques such as ANNs leads to more accurate estimations of
probability distribution parameters for downtime lengths and their associated factors.
Experimental results show that, in certain scenarios, the proposed model achieves per-
formance improvements of up to 30% when compared to established approaches in the
literature.
The implications of this research extend beyond academia to industry practitioners
and decision-makers. By accurately modeling downtime lengths, the proposed frame-
work enables proactive maintenance scheduling and resource allocation, ultimately
enhancing effectiveness and productivity in industrial operations. However, several
research directions remain open for exploration. The robustness and scalability of the
proposed framework across diverse industrial contexts should be investigated. Addition-
ally, introducing real-time data stream processing with predictive analytics techniques
could enable real-time failure prediction and estimation of accompanying downtime
lengths.
Supplementary Information
The online version contains supplementary material available at [Link]
Supplementary Material 1
Supplementary Material 2
Author contributions
V. K. and K. P. conceived and designed the study and developed the methodology. V. K. supervised data collection and
provided technical expertise in paper manufacturing processes. K. P. and A. M. performed statistical data analysis and
interpreted the results. K. P. implemented machine learning algorithms. A. M. implemented statistical techniques. V. K.,
K. P. and A. M. wrote the main manuscript text. V. K., K.P. and S. K. prepared figures. S. K. contributed to the development
of the methodology, assisted in experimental design and critically reviewed the manuscript. I.M. contributed to the
conceptualization of the research, reviewed the manuscript for accuracy and completeness. V. B. provided administrative
and logistical support and reviewed the manuscript for accuracy and completeness.
Declarations
Competing interests
The authors declare no competing interests.
References
1. Shah R, Ward P. Lean manufacturing: context, practice bundles, and performance. J Oper Manag. 2003;21(2):129–49.
2. Okpala C, Anozie S. Overall equipment effectiveness and the six big losses in total productive maintenance. J Sci
Eng Res. 2018;5(4):156–64.
3. Alsyouf I. The role of maintenance in improving companies’ productivity and profitability. Int J Prod Econ.
2007;105(1):70–8.
4. Bokrantz J, Skoogh A, Berlin C, Stahre J. Maintenance in digitalised manufacturing: Delphi-based scenarios for 2030.
Int J Prod Econ. 2017;191:154–69.
5. Banković M, Filipović V, Graovac J, Hadži-Purić J, Hurson AR, Kartelj A, Kovaččević J, Korolija N, Kotlar M, Krdžavac
NB, Marić F, Malkov S, Milutinović V, Mitić N, Mišković S, Nikolić M, Pavlović-Lažetić G, Simić D, Stojanović Djurdjević
S, Vujičić Stanković S, Vujošević Janičić M, Živković M. Chapter one—teaching graduate students how to review
research articles and respond to reviewer comments. Advances in Computers, vol. 116. Elsevier; 2020. p. 1–63.
6. Blagojević V, Bojić D, Bojović M, Cvetanović M, Djordjević J, Djurdjević D, Furlan B, Gajin S, Jovanović Z, Milićev D,
Milutinović V, Nikolić B, Protić J, Punt M, Radivojević Z, Stanisavljević Ž, Stojanović S, Tartalja I, Tomašević M, Vuletić P.
A systematic approach to generation of new ideas for PhD research in computing. In: Creativity in computing and
dataflow supercomputing. Advances in computers, vol 104; 2017. p. 1–31
7. Koković V, Mačužić I, Todorović P. Application of performance-based maintenance in Lean production systems. In:
Proceedings of the 19th international scientific conference on industrial systems (IS’23), VP1.1. 7–10241, Novi Sad,
Serbia; 2023.
8. SKF: rolling bearings in paper machines A handbook for paper machine designers, operators, and maintenance
staff; 2016.
9. SKF: Acceleration enveloping in paper machines: an approach to extracting very low frequency impact signals.
Application Note CM3024 EN; 2011.
10. Nakajima S. TPM development program: implementing total productive maintenance; 1989.
11. SMRP: Society of Maintenance and Reliability Professionals (SMRP) Best practices 6th Edition; 2020.
12. Surucu O, Gadsden SA, Yawney J. Condition monitoring using machine learning: a review of theory, applications,
and recent advances. Expert Syst Appl. 2023;221: 119738.
13. Fernandes M, Corchado JM, Marreiros G. Machine learning techniques applied to mechanical fault diagnosis and
fault prognosis in the context of real industrial manufacturing use-cases: a systematic literature review. Appl Intell.
2022;52(12):14246–80.
14. Shahin M, Chen FF, Hosseinzadeh A, Zand N. Using machine learning and deep learning algorithms for downtime
minimization in manufacturing systems: an early failure detection diagnostic service. Int J Adv Manuf Technol.
2023;128(9–10):3857–83.
15. El-Mazgualdi C, Masrour T, El-Hassani I, Khdoudi A. Machine learning for KPIS prediction: a case study of the overall
equipment effectiveness within the automotive industry. Soft Comput. 2021;25(4):2891–909.
16. Brunelli L, Masiero C, Tosato D, Beghi A, Susto GA. Deep learning-based production forecasting in manufacturing: a
packaging equipment case study. Procedia Manuf. 2019;38:248–55.
17. Gomilanović M, Stanić N, Milijanović D, Stepanović S, Milijanović A. Predicting the availability of continuous mining
systems using LSTM neural network. Adv Mech Eng. 2022;14(2):16878132221081584.
18. Reyneri L, Colla V, Vannucci M. Estimate of a probability density function through neural networks. In: Advances in
computational intelligence: proceedings of the 11th international work-conference on artificial neural networks
(IWANN 2011), Part I 11. Springer; 2011. p. 57–64.
19. Liu Q, Xu J, Jiang R, Wong WH. Density estimation using deep generative neural networks. Proc Natl Acad Sci.
2021;118(15):2101344118.
20. Chen CH, Song F, Hwang FJ, Wu L. A probability density function generator based on neural networks. Phys A Stat
Mech Appl. 2020;541: 123344.
21. Barnett V, Lewis T. Outliers in statistical data; 1994.
22. Smirnov N. Table for estimating the goodness of fit of empirical distributions. Ann Math Stat. 1948;19(2):279–81.
23. Wang C, Zeng B, Shao J. Application of bootstrap method in Kolmogorov-Smirnov test. In: International conference
on quality, reliability, risk, maintenance, and safety engineering, Xi’an, China; 2011. p. 287–291.
24. Morgenstern D. Economic prediction and statistical distribution. J Am Stat Assoc. 1952;47(259):792–8.
25. Gumbel EJ. Bivariate exponential distributions. J Am Stat Assoc. 1952;47(259):796–7.
26. Farlie DJG. The performance of some variance-covariance type estimates. Biometrika. 1960;47(3/4):307–23.
27. Vugrin K, Swiler L, Roberts R, Stucky-Mack N, Sullivan S. Confidence region estimation techniques for nonlinear
regression in groundwater flow: three case studies. Water Resour Res. 2007. [Link]
28. Fisher R. Statistical methods for research workers. Statistical methods for research workers; 1936.
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
1. use such content for the purpose of providing other users with access on a regular or large scale basis or as a means to circumvent access
control;
2. use such content where to do so would be considered a criminal or statutory offence in any jurisdiction, or gives rise to civil liability, or is
otherwise unlawful;
3. falsely or misleadingly imply or suggest endorsement, approval , sponsorship, or association unless explicitly agreed to by Springer Nature in
writing;
4. use bots or other automated methods to access the content or redirect messages
5. override any security feature or exclusionary protocol; or
6. share the content in order to create substitute for Springer Nature products or services or a systematic database of Springer Nature journal
content.
In line with the restriction against commercial use, Springer Nature does not permit the creation of a product or service that creates revenue,
royalties, rent or income from our content or its inclusion as part of a paid for service or for other commercial gain. Springer Nature journal
content cannot be used for inter-library loans and librarians may not upload Springer Nature journal content on a large scale into their, or any
other, institutional repository.
These terms of use are reviewed regularly and may be amended at any time. Springer Nature is not obligated to publish any information or
content on this website and may remove it or features or functionality at our sole discretion, at any time with or without notice. Springer Nature
may revoke this licence to you at any time and remove access to any copies of the Springer Nature journal content which have been saved.
To the fullest extent permitted by law, Springer Nature makes no warranties, representations or guarantees to Users, either express or implied
with respect to the Springer nature journal content and all parties disclaim and waive any implied warranties or warranties imposed by law,
including merchantability or fitness for any particular purpose.
Please note that these rights do not automatically extend to content, data or other material published by Springer Nature that may be licensed
from third parties.
If you would like to use or distribute our Springer Nature journal content to a wider audience or on a regular basis or in any other manner not
expressly permitted by these Terms, please contact Springer Nature at
onlineservice@[Link]
Using historical performance data is crucial for forecasting future KPIs as it enables companies to understand trends and patterns in equipment behavior. This insight allows for more informed maintenance decisions, improving the alignment of maintenance strategies with performance-based goals. Forecasting helps identify the realistic achievements of improvement programs and enables structured incentives aligned with KPI results, thus enhancing overall maintenance effectiveness .
The use of sigmoid activation functions in hidden layers facilitates learning complex, non-linear relationships in the data, essential for accurately capturing downtime distributions. The softplus activation in the output layer is well-suited for modeling distribution parameters such as those in exponential and gamma distributions, as it produces only non-negative outputs, ensuring alignment with the positive nature of downtime lengths. This combination enhances the model's capacity to learn and predict downtime distributions better .
Big data analytics aids in optimizing maintenance schedules, prioritizing critical tasks, and efficiently allocating resources. It enables the prediction of key performance indicators (KPIs) by analyzing historical performance data, which helps in identifying potential issues before they occur and enhances overall performance. This analytical process involves data collection from sensors, detailed data analysis, and generating alerts for future failures, thereby improving maintenance planning and execution .
Unsupervised learning methods, such as those using ANNs for estimating distribution parameters, enable modeling downtime lengths without the need for labeled datasets, which are often large and unlabelled in industrial settings. This approach enhances model generalization and adaptability, allowing for improved prediction accuracy across varying operational conditions, and can be expanded by integrating diverse variables like machine vibrations .
Discretizing data into intervals can introduce a bias by flattening the data's richness and potentially losing critical variability and nuances. In the context of downtime length modeling, it may lead to skewed model predictions due to improperly captured variance within each interval. The choice of interval size can significantly influence model performance, as seen in discrepancies when comparing models trained on discrete and continuous domains. This impacts the precision of predicted distributions and affects decision-making in maintenance planning .
The increase in fiber-based packaging demand could lead to higher production levels in paper manufacturing, necessitating more efficient operations and maintenance strategies. Companies may adopt more advanced predictive maintenance and condition monitoring techniques, such as vibration analysis, to ensure high machine uptime and reliability to meet market demands. The emphasis might shift from cost-centric maintenance to value-driven performance-based strategies, aligning with broader sustainability trends by offering alternatives to plastic packaging .
The prevalence of zero downtime values in datasets can bias ANN models towards predicting non-failure scenarios. To address this issue, the training focuses more on precisely estimating longer downtime lengths, which are crucial as they represent actual failure events impacting productivity. Additionally, weights in ANNs are regularized using the L2-norm to mitigate overfitting and counteract the impact of the zero values bias .
ANN models outperform baseline statistical models for both mechanical and electrical failure data, as indicated by lower negative log-likelihood (NLL) values. For mechanical failures, the ANN model achieved an NLL of 389.69 compared to the baseline's 394.34. For electrical failures, the ANN's NLL was 595.48, significantly better than the baseline's 709.98. The comparative model from [17] showed even higher NLLs, indicating lesser adaptability for both failure types .
Utilizing ANNs to estimate parameters of probability distributions, rather than directly modeling probability density function (PDF) values, leads to significant performance improvements by up to 30% compared to alternative methods. This approach also facilitates unsupervised training, making it suitable for large repositories of unlabelled data and allowing for potential expansion by incorporating additional input variables such as machine vibration velocity measurements .
Vibration analysis acts as an early warning system for potential machine failures, allowing time for maintenance planning to prevent catastrophic failures. It provides insights into the condition of critical components, such as bearings and gears, and is instrumental in identifying adverse conditions related to machine reliability in the harsh environments of paper machines .