Cost-Based FMEA for Improved Reliability
Cost-Based FMEA for Improved Reliability
[Link]/locate/aei
Abstract
Failure Modes and Effects Analysis (FMEA) is a design tool that mitigates risks during the design phase before they occur. Although many
industries use the current FMEA technique, it has many limitations and problems. Risk is measured in terms of Risk Priority Number (RPN)
that is a product of occurrence, severity, and detection difficulty. Measuring severity and detection difficulty is very subjective and with no
universal scale. RPN is also a product of ordinal variables, which is not meaningful as a proper measure. This paper addresses these
shortcomings and introduces a new methodology, Life Cost-Based FMEA, which measures risk in terms of cost. Life Cost-Based FMEA is
useful for comparing and selecting design alternatives that can reduce the overall life cycle cost of a particular system. Next, a Monte Carlo
simulation is applied to the Cost-Based FMEA to account for the uncertainties in: detection time, fixing time, occurrence, delay time, down
time, and model complex scenarios. A case study of a large scale particle accelerator shows the advantages of the proposed approach in
predicting life cycle failure cost, measuring risk and planning preventive, scheduled maintenance and ultimately improving up-time.
q 2004 Elsevier Ltd. All rights reserved.
Keywords: FMEA; Life cost-based FMEA; Failure cost; Reliability; Availability; Empirical data
hotels, restaurants, and movies. Ordinal values preserve The rating is scaled from 1 to 10 for each category. The
rank but the distance between the values cannot be occurrence is related to the probability of the failure mode
measured since a distance function does not exist. Thus, and cause. Occurrence ratings have been standardized by
the RPN, which is a product of three independent variables, many electronics and automotive industries [11] over the
is not meaningful. last few years. A ‘10’ on the occurrence table corresponds to
a failure happening with every other part. A ‘1’ corresponds
1.3. Related research to one failure in a million parts.
The severity index measures the seriousness of the
Recent FMEA research has been focused on improving effects of a failure mode. Thus, a severity index is assigned
traditional FMEA limitations by using different measure- to the end effect of a failure. A ‘1’ on the severity index
ment schemes, considering multiple failure scenarios, corresponds to a failure that does not affect anything, a ‘5’
and incorporating sensitivity analysis. Selected samples corresponds to a performance loss, a ‘7’ corresponds to
of recent research in FMEA include the following: machine shut down, and a ‘10’ corresponds to a life-
threatening failure.
† Tracing causal chains and their probabilities using The detection index is generated on the basis of the
Bayesian Networks [2]. likelihood of detection by relevant design reviews, testing,
† Using a Petri net to analyze multiple failure effects [5]. and quality control measures. A ‘1’ on the detection index
† Identifying and prioritize the process part of potential corresponds to a failure mode that is almost certain to be
problems that have the most financial impact on an detected and a ‘10’ corresponds to a failure that is almost
operation [6]. impossible to detect. Taking the product of these three
† Using probability of a certain failure and the probability indices (occurrence, severity, and detection) generates the
that this failure will not be detected to obtain expected RPN. The RPN represents the risk associated to each failure
failure cost [7]. mode.
† Using RPN on a logarithmic scale [8].
† Applying Monte Carlo simulation on RPN numbers [9].
2.2. Life cost-based FMEA
† Using occurrence and severity as a risk measure for
FMECA [10].
To resolve the ambiguity of measuring detection
These new FMEA approaches have addressed some of difficulty and the irrational logic of multiplying three
the problems mentioned in the previous section but not yet ordinal indices, a new methodology was created to
adequately addressed how to: (1) determine failure cost, overcome shortcomings, Life Cost-Based FMEA. Life
(2) address sensitivity analysis, and (3) resolved confusion Cost-Based FMEA measures failure/risk in terms of cost
with detection. [14]. Cost is a universal language that can be easily
The investigation presented in this paper builds upon understood in terms of severity among engineers and others.
earlier research [3], which is based on scenario-based FMEA Thus, failure cost can be estimated using the following
to weigh the expected life cost of failure during the early part simplest form:
of design. Shortcomings of traditional FMEA will be resolved X
n
through the introduction of cost as a measure of risk in this Expected failure cost Z p i ci (1)
iZ1
paper. Failures may occur at any stage of the product
development life cycle: design, manufacturing, installation,
p probability of a particular failure occurring
and operation. Failure cost becomes greater as the origin and
c cost associated with that particular failure
detection stages of a failure become further apart in time.
The case study for this paper was done in conjunction
Table 1 shows a Life Cost-based FMEA table created
with and supported by research and development being
for the methodology. The frequency value can be either
performed at the Stanford Linear Accelerator Center
the probability or the frequency of occurrence. Failures
(SLAC) for the Next Linear Collider (NLC). All of the
that originate in design, manufacturing, and installation
quantitative estimates in this work should be considered as
are assumed to be one-time event failures and the
illustrative only, and do not reflect what the actual costs
probability of occurrence is assigned to the frequency
might be at some time in the future.
variable. Failures that originate in assembly and opera-
tions reoccur during the life-time of the system thus
2. FMEA methodologies frequency of failure during a 1 year period is assigned to
this variable. Re-occurring variable indicate whether the
2.1. Traditional RPN failure is a one time event or reoccurs over the life time
of the product.
A traditional FMEA uses RPN to assess risk in three Failure origin indicates when the failure has been
categories: Occurrence (O), Severity (S), and Detection (D). initially introduced. Detection phase indicates the stage at
S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188 181
Opportu-
nity cost
112,500
125,000
250,000
62,500
($)
Material
cost ($)
3000
5000
4500
15
180
115,200
38,400
1280
cost ($)
Output
Labor
50
50
50
cost
($)
which the failure has been realized. Fig. 1 shows the four
Quan-
the failure is a design error. Due to design error, the part has
to be redesigned, remanufactured, and reinstalled. There
Delay
time
Failures may occur at any stage of the life cycle and can
4
0.5
0.5
30
30
1
Oper
Oper
TR
Oper
Oper
Oper
Inst
turned off
turned off
turned off
Effect of
Magnet
Magnet
Magnet
Magnet
failure
sprayed on
Too many
passage is
of failure
loads on
blocked
to coil
circuit
Water
Water
overheating
trip due to
type of material being used. These mistakes can be detected 3. Applying empirical data on life cost based FMEA
during parts inspection or in the subsequent stages.
Examples of failures during the installation stage are In electrical power plant and chemical process industries,
following incorrect installation procedure, applying too LCC analysis is more closely linked to system availability
much or too little force on to tools when tightening analysis than other industries, because production regularity
fasteners, damaging the part, etc. Labor cost can be derived is one of the biggest concerns for plant owners. LCC
with the time information obtained in the cost-based FMEA analysis in plant industries tends to focus on prediction of
table using the following equation: the unavailability of the total system due to component
failures, maintenance and emergency shutdowns.
Labor cost Z occurrence !f½detection time The availability of a repairable component is approxi-
!labor rate !no: of operators mated as expressed in Eq. (5), if after each repair ‘as good as
new’ is assumed [12].
C ½fixing time !labore rate
† Availability. Average probability that an item will
!no: of operators C ½delay time perform its required function under given conditions at
time.
!labor rate !no: of operatorsg (2)
MTTF MTTF
Availability ðAÞ Z Z (5)
Component replacement due to failure is considered as MTTF C MTTR MTBF
material cost. Material cost is obtained using the following
equation: † MTBF (Mean Time Between Failures). MTBF is a basic
Material cost Z occurrence !cost of part (3) measure of reliability for repairable items. It can be
described as the number of hours that pass before a
Opportunity cost is the cost that incurs when a failure component, assembly, or system fails. It is a commonly
inhibits the main function of the system and prevents any used variable in reliability and maintainability analyses.
creation of value. Opportunity cost is the cost incurred when † MTTR (Mean Time to Repair). MTTR is the average time
a failure inhibits the main function of the system and required to perform corrective maintenance on all of the
prevents any creation of value. Opportunity cost is removable items in a product or system. This kind of
calculated using the following equation: maintainability prediction analyzes how long repairs and
Opportunity cost maintenance tasks will take in the event of a system failure.
† MTTF (Mean Time To Failure). MTTF is a basic measure
Z down time !hourly opportunity cost (4) of reliability for non-repairable systems. It is the mean time
expected until the first failure of a piece of equipment.
where MTTF is a statistical value and is meant to be the mean over
Down time Z fdetection time C fixing time C delay timeg a long period of time and large number of units.
Table 4
Run time of water cooled electromagnet
Magnets System
The third column shows the number of water-cooled the following equation:
electromagnets for that particular line. The fifth column is AMSys Z ASM !AWM Z 0:9987 !0:9549 Z 0:9536 (11)
the product of run hour and the number of magnets: magnet
hours. The sixth column indicates the number of failures
identified during that particular period. The MTBF in the ASM availability of solid wire magnet
seventh column is a result of magnet hours divided by the AWM availability of water-cooled magnet
number of failures. The eighth column indicates the total
repair time for those failures in that period and the ninth Thus, this would fall short of the 97.5% availability goal
column is MTTR. Based on these numbers, the availability if the design of the new magnets does not eliminate the
of any one magnet in a beamline can be calculated. root cause of the observed failures. A summary of the
The average availability of one water-cooled magnet at availability is shown in Table 5.
SLAC is found to be 0.9999907. This example predicts the overall failure for the NLC,
but one can predict failures for particular types of failure
The availability of the NLC’s electromagnet subsystem
(insulation, water leak, water blockage, mechanical, or
can be estimated using Eq. (7). Assuming the reliability of
human error) using the same methodology.
each individual magnet is 0.9999907 the availability for
4965 water-cooled electromagnets would be 0.9548. 4.2. Power supply
However, this is lower than the target value of 97.5% for
the magnet subsystem. Therefore, the magnet designers The power supplies that provide the electric current to
know they must improve the reliability of the magnets they the electromagnets can be categorized into two main
design for NLC over the SLAC magnets. Given 6489 h of
operation time per year, the expected downtime of the NLC Table 5
due to electromagnet failure is 292 h/yr. Since the average Predicted availability of electromagnets for NLC
MTTR is 10 h, we can estimate the number of failures for a Type Solid wire Water-cooled
given year to be 29 occurrences. No. of Magnets 2202 4965
Availability of solid wire magnets can be calculated in Availability 0.9987 0.9548
the same manner. The expected number of failures for solid Expected downtime 8.3 292
(h/yr)
wire magnets in the NLC is twice a year. The overall
Occurrence per year 1.9 29.2
availability of the NLC magnet system is obtained using
S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188 185
Table 6
Downtime of accelerator due to power supply failure
Units: hour.
Fig. 2. Electromagnet system.
Fig. 3. Monte Carlo simulation of labor and material cost for electromagnet.
Table 8
Predicted life cycle failure cost of electromagnets for the NLC for 30 years
Units: million.
are considered, and $50K when the cost of building the NLC large power supplies is still quite high because the power
is amortized over a 30-year period in addition to the labor supply electric boards have to be replaced regardless of the
and energy cost. Thus, the overall opportunity cost was shutdown of the accelerator.
calculated for all three values. The magnet system requires the electromagnets and
A Monte Carlo simulation is applied to the Life Cost- power subsystem to both be working. Thus, the life cycle
Based FMEA to consider the sensitivity of variables failure cost of the subsystem is the sum of electromagnet
associated to failure cost: frequency, detection time, fixing and power supply failure cost as shown in Table 10. The
time, delay time, and parts cost. Fig. 3 shows result of the actual labor and material cost is a small fraction of what the
simulation for labor and material costs for the different total opportunity cost might be, even using the lowest
confidence levels. A 30 year predicted failure cost for the opportunity cost per hour, $10K/h.
electromagnet is summarized in Table 8. As shown in the
table, opportunity cost can be 30–150 times greater than
the labor and material cost. 5. Discussion
The estimated failure cost for the system of power
supplies is summarized in Table 9. As predicted in Table 7, As derived in Section 4, availability for the electromag-
the availability of large power supplies is pretty low, 0.938. net system falls short of the target goal of 97.5%. To
Thus, redundancy is assumed for the large power supplies to increase the availability of the water-cooled magnets for
meet the availability goal. Material and labor failure cost for the NLC, two measures can be taken: reduce MTTR or
Table 9
Life cycle failure cost of power supply of 30 years Table 10
Life cycle failure cost of electromagnet system
Small Large Total
Failure cost
Labor cost $0.39 $1.90 $2.29
Material cost $0.92 $7.20 $8.12 Labor cost $4.2M
Sub total $1.31 $9.10 $10.41 Material cost $9.3M
Opportunity cost $10K $23 $6 $29 Sub total $13.5M
$25K $59 $15 $74 Opportunity cost $10K $126.8M
$50K $117 $30 $147 $25K $318.2M
$50K $635M
Units: million.
S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188 187
Units: million.
188 S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188
manufacturing, installation, and operation. Designers can Cherrill Spencer and John Cornuelle for their valuable time
readily incorporate the changes in the model to estimate an spent on this research, and other SLAC staff members who
improved life cycle cost. The root causes directs designers provided us with useful technical information.
to focus their efforts on problem systems, components, and
processes.
Complex systems usually have set target availability.
One means to achieve the target is to increase all subsystems References
reliabilities. However, guaranteeing higher reliability often
incurs cost increases. Another solution is to schedule [1] Stamatis DH. Failure mode and effect analysis. Milwaukee, WI: ASQ
preventive maintenance. Our proposed methodology maps Quality Press; 1995.
[2] Lee B. Using Bayes belief networks in industrial FMEA modeling and
allow comparisons of different availability enhancement
analysis. Proceedings of International Symposium on Product Quality
measures and trace analysis in terms of cost, a widely and Integrity, Philadelphia, PA; 2000.
accepted measure of risk. [3] Kmenta S, Ishii K, Scenario-based FMEA: a life cycle cost
The authors agree that extracting relevant knowledge perspective. Proceedings of ASME Design Engineering Technical
from pre-existing data (CATER system) is hard work Conference, Baltimore, MD; 2000.
[4] Palady P. Failure modes and effects analysis; predicting and
because it collects data without the purpose of improving
preventing problems before they occur. Florida: PT Publications;
reliability or serviceability. It is evident that reliability and 1995.
serviceability should be considered when maintenance [5] He D, Adamyan A. An impact analysis methodology for design of
management systems are put into place. Many times the products and processes for reliability and quality. Proceedings of
management system records data in a way that it cannot be ASME Design Engineering Technical Conference, Pittsburgh, PA;
2001.
used effectively. One could apply data mining techniques to
[6] Tarum CD. FMERA—failure modes, effects, and (financial) risk
improve feedback from experience to extract relevant data analysis. SAE World Congress, Detroit, MI; 2001.
from tons of data that have been recorded [17]. The example [7] Gilchrist W. Modeling failure modes and effects analysis. Int J Quality
presented in this paper used a semi-manual sorting Reliab Manage 1992;10:16–23.
technique to extract the relevant data. [8] Ben-Daya M, Raouf A. A revised failure mode and effects analysis
Life Cost-Based FMEA can also provide a fair model. Int J Quality Reliab Manage 1996;13:43–7.
[9] Bevilacqua M, Braglia M, Gabbrielli R. Monte Carlo simulation
comparison between competing designs of subsystems. approach for modified FMECA in a power plant. Quality Reliab Eng
The case study presented in this paper considered only Int 2000;16:313–24.
the currently used magnet technology. The proposed [10] MIL-STD-1629A, BS 5760 Part 5.
methodology may not simply extrapolate to new and/or [11] SAE ARP-4293: Life cycle cost—techniques and applications.
unproven technology, because empirical or expert knowl- [12] Birolini A. Quality and reliability of technical systems, 2nd ed. Berlin:
Springer; 1997.
edge may not be available. Thus, future research lies in [13] Sass R, Shoaee H. CATER: an online problem tracking facility for
estimating uncertainty variables (e.g. frequency, detection SLC. Proceedings of Particle Accelerator Conference, Washington
time, fixing time, and delay time) using available com- DC, 1993.
ponent data, and extrapolating them to higher subsystem [14] Rhee S, Ishii K. Life cost based FMEA incorporating data uncertainty.
levels. Hybrid use of empirical and analytical data will Proceedings of ASME Design Engineering Technical Conference,
Montreal, Canada; 2002.
present significant new challenges.
[15] Mettas A. Reliability allocation and optimization for complex
systems. Proceedings of IEEE Reliability and Maintainability
Symposium, Philadelphia, PA; 2001.
Acknowledgements [16] Gershenson J, Ishii K. Design for serviceability. In: Kusiak A, editor.
Concurrent engineering: theory and practice. New York: Wiley; 1992.
p. 19–39.
This research has been supported by the Department of
[17] Manago M, Auriol E. Using data mining to improve feedback from
Energy contract, DE-AC03-76SF00515. The authors would experience for equipment in the manufacturing and transport
like to thank the Stanford Linear Accelerator Center for industries.: Institute for Operations Research and the Management
providing the opportunity for this research, and especially Scienece; 1996.