0% found this document useful (0 votes)
16 views10 pages

Cost-Based FMEA for Improved Reliability

Uploaded by

matuidi512
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
16 views10 pages

Cost-Based FMEA for Improved Reliability

Uploaded by

matuidi512
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Advanced Engineering Informatics 17 (2003) 179–188

[Link]/locate/aei

Using cost based FMEA to enhance reliability and serviceability


Seung J. Rhee, Kosuke Ishii*
Department of Mechanical Engineering, Design Division, Stanford University, Stanford CA 94305, USA

Abstract

Failure Modes and Effects Analysis (FMEA) is a design tool that mitigates risks during the design phase before they occur. Although many
industries use the current FMEA technique, it has many limitations and problems. Risk is measured in terms of Risk Priority Number (RPN)
that is a product of occurrence, severity, and detection difficulty. Measuring severity and detection difficulty is very subjective and with no
universal scale. RPN is also a product of ordinal variables, which is not meaningful as a proper measure. This paper addresses these
shortcomings and introduces a new methodology, Life Cost-Based FMEA, which measures risk in terms of cost. Life Cost-Based FMEA is
useful for comparing and selecting design alternatives that can reduce the overall life cycle cost of a particular system. Next, a Monte Carlo
simulation is applied to the Cost-Based FMEA to account for the uncertainties in: detection time, fixing time, occurrence, delay time, down
time, and model complex scenarios. A case study of a large scale particle accelerator shows the advantages of the proposed approach in
predicting life cycle failure cost, measuring risk and planning preventive, scheduled maintenance and ultimately improving up-time.
q 2004 Elsevier Ltd. All rights reserved.

Keywords: FMEA; Life cost-based FMEA; Failure cost; Reliability; Availability; Empirical data

1. Introduction a couple of columns to describe the entire fault chain [2],


inhibiting the understanding of the true cause of failures.
1.1. Introduction to FMEA Thus, a more thorough analysis, such as scenario based
FMEA [3], can be used to understand all intermediate
Failure Modes and Effects Analysis (FMEA) is a tool effects between the initiating cause to end effects.
widely used in the automotive, aerospace, and electronics
industries to identify, prioritize, and eliminate known
1.2. Shortcomings of traditional FMEA
potential failures, problems, and errors from systems
under design before the product is released [1]. Several
industrial FMEA standards such as the Society of One definition of detection (D) difficulty is how well the
Automotive Engineers, US Military of Defense, and organization controls the development process. Another
Automotive Industry Action Group employ the Risk Priority definition relates to the detectability of failure on the product
Number (RPN) to measure risk and severity of failures. RPN is in the hands of the customer. The former asks ‘What is
is a product of three indices: Occurrence (O), Severity (S), the chance of catching the problem before we give it to the
and Detection (D). Design engineers typically analyze the customer?’ The latter asks ‘What is the chance of
‘root cause’ and ‘end-effects’ of potential failures in the sub- the customer catching the problem before the problem
system or component. The analysis is organized around results in a catastrophic failure?’ [4]. These definitions
failure modes, which link the cause and effect of failures. confuse the FMEA users when one tries to determine
Traditional FMEA sheets limit failure representation to only detection difficulty. Are we trying to measure how easy it is
to detect where a failure has occurred or when it has
occurred? On the other hand, are we trying to measure how
* Corresponding author. Tel.: C1-650-725-1840; fax: C1-650-723-
7349.
easy or difficult it is to prevent failures?
E-mail addresses: rhees@[Link] (S.J. Rhee), ishii@ The three indices used for RPN are ordinal scale
[Link] (K. Ishii). variables that are used to rank-order industries such as,
1474-0346/$ - see front matter q 2004 Elsevier Ltd. All rights reserved.
doi:10.1016/[Link].2004.07.002
180 S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188

hotels, restaurants, and movies. Ordinal values preserve The rating is scaled from 1 to 10 for each category. The
rank but the distance between the values cannot be occurrence is related to the probability of the failure mode
measured since a distance function does not exist. Thus, and cause. Occurrence ratings have been standardized by
the RPN, which is a product of three independent variables, many electronics and automotive industries [11] over the
is not meaningful. last few years. A ‘10’ on the occurrence table corresponds to
a failure happening with every other part. A ‘1’ corresponds
1.3. Related research to one failure in a million parts.
The severity index measures the seriousness of the
Recent FMEA research has been focused on improving effects of a failure mode. Thus, a severity index is assigned
traditional FMEA limitations by using different measure- to the end effect of a failure. A ‘1’ on the severity index
ment schemes, considering multiple failure scenarios, corresponds to a failure that does not affect anything, a ‘5’
and incorporating sensitivity analysis. Selected samples corresponds to a performance loss, a ‘7’ corresponds to
of recent research in FMEA include the following: machine shut down, and a ‘10’ corresponds to a life-
threatening failure.
† Tracing causal chains and their probabilities using The detection index is generated on the basis of the
Bayesian Networks [2]. likelihood of detection by relevant design reviews, testing,
† Using a Petri net to analyze multiple failure effects [5]. and quality control measures. A ‘1’ on the detection index
† Identifying and prioritize the process part of potential corresponds to a failure mode that is almost certain to be
problems that have the most financial impact on an detected and a ‘10’ corresponds to a failure that is almost
operation [6]. impossible to detect. Taking the product of these three
† Using probability of a certain failure and the probability indices (occurrence, severity, and detection) generates the
that this failure will not be detected to obtain expected RPN. The RPN represents the risk associated to each failure
failure cost [7]. mode.
† Using RPN on a logarithmic scale [8].
† Applying Monte Carlo simulation on RPN numbers [9].
2.2. Life cost-based FMEA
† Using occurrence and severity as a risk measure for
FMECA [10].
To resolve the ambiguity of measuring detection
These new FMEA approaches have addressed some of difficulty and the irrational logic of multiplying three
the problems mentioned in the previous section but not yet ordinal indices, a new methodology was created to
adequately addressed how to: (1) determine failure cost, overcome shortcomings, Life Cost-Based FMEA. Life
(2) address sensitivity analysis, and (3) resolved confusion Cost-Based FMEA measures failure/risk in terms of cost
with detection. [14]. Cost is a universal language that can be easily
The investigation presented in this paper builds upon understood in terms of severity among engineers and others.
earlier research [3], which is based on scenario-based FMEA Thus, failure cost can be estimated using the following
to weigh the expected life cost of failure during the early part simplest form:
of design. Shortcomings of traditional FMEA will be resolved X
n

through the introduction of cost as a measure of risk in this Expected failure cost Z p i ci (1)
iZ1
paper. Failures may occur at any stage of the product
development life cycle: design, manufacturing, installation,
p probability of a particular failure occurring
and operation. Failure cost becomes greater as the origin and
c cost associated with that particular failure
detection stages of a failure become further apart in time.
The case study for this paper was done in conjunction
Table 1 shows a Life Cost-based FMEA table created
with and supported by research and development being
for the methodology. The frequency value can be either
performed at the Stanford Linear Accelerator Center
the probability or the frequency of occurrence. Failures
(SLAC) for the Next Linear Collider (NLC). All of the
that originate in design, manufacturing, and installation
quantitative estimates in this work should be considered as
are assumed to be one-time event failures and the
illustrative only, and do not reflect what the actual costs
probability of occurrence is assigned to the frequency
might be at some time in the future.
variable. Failures that originate in assembly and opera-
tions reoccur during the life-time of the system thus
2. FMEA methodologies frequency of failure during a 1 year period is assigned to
this variable. Re-occurring variable indicate whether the
2.1. Traditional RPN failure is a one time event or reoccurs over the life time
of the product.
A traditional FMEA uses RPN to assess risk in three Failure origin indicates when the failure has been
categories: Occurrence (O), Severity (S), and Detection (D). initially introduced. Detection phase indicates the stage at
S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188 181

Opportu-
nity cost

112,500

125,000

250,000
62,500
($)
Material
cost ($)

3000

5000

4500
15
180

115,200
38,400

1280
cost ($)
Output

Labor

Fig. 1. Initial origin and detection stages of failure.


1250
Parts

50

50

50
cost
($)

which the failure has been realized. Fig. 1 shows the four
Quan-

different initiating stages (design, manufacture, installation,


tity

operation) and the failure detection stages (design review,


inspection, testing, operation), and an example where failure
Loss
time

is detected during operations stage and the initial origin of


10
5

the failure is a design error. Due to design error, the part has
to be redesigned, remanufactured, and reinstalled. There
Delay
time

may be some delay between each activity also. Recovery


0

time is the time the system is inoperable due to failure.


Recovery time is associated with lost opportunity.
Fixing
time

Failures may occur at any stage of the life cycle and can
4

be detected either during the same stage or during


subsequent stages. The failure cost is minimal when the
Detection

origin and detection occur during the same stage. The


time

0.5

0.5

failure cost increases as the origin and detection stages


1

become further apart.


Frequency

Failure cost has three major components: labor cost,


material cost, and opportunity cost. Labor cost and
0.001

opportunity cost can be measured in terms of time and can


2

be further broken up into four different stages: detection


Re-occur-

time, fixing time, delay time, and recovery time.


ring

† Detection time. Time to realize and identify a certain type


30

30

30
1

of failure has occurred and diagnoses the exact location.


Detection

† Fixing time. Time to fix the problem. The actual fixing


phase

time for each individual component.


Oper

Oper

Oper
TR

† Delay time. Time incurred for non-value activity such as


waiting for response, set up time, and mailing/shipping
time.
Origin
Input

Oper

Oper

Oper
Inst

† Recovery time. Time to have the system up and running


to its original state. Only applies to failures that happen
during the operations stage.
turned off

turned off

turned off

turned off
Effect of

Magnet

Magnet

Magnet

Magnet
failure

Examples of failures during the design stage are incorrect


design calculations, incorrect number in prints, and incorrect
Root cause

sprayed on
Too many

passage is

material selection. These mistakes may be detected during


Damaged
Life cost-based FMEA table

of failure

loads on

blocked

to coil
circuit
Water

Water

design reviews before the components are manufactured or in


coil

the subsequent stages. Failure cost is minimized when failure


and detection occurs in the same stage. The most expensive
Thermal switch

failure is which originates during the design stage and does


Failure mode

overheating
trip due to

not get detected until operation (customer detection).


Table 1

Examples of failures during the manufacturing stage are


operator’s mistake, bad material calibration, and incorrect
182 S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188

type of material being used. These mistakes can be detected 3. Applying empirical data on life cost based FMEA
during parts inspection or in the subsequent stages.
Examples of failures during the installation stage are In electrical power plant and chemical process industries,
following incorrect installation procedure, applying too LCC analysis is more closely linked to system availability
much or too little force on to tools when tightening analysis than other industries, because production regularity
fasteners, damaging the part, etc. Labor cost can be derived is one of the biggest concerns for plant owners. LCC
with the time information obtained in the cost-based FMEA analysis in plant industries tends to focus on prediction of
table using the following equation: the unavailability of the total system due to component
failures, maintenance and emergency shutdowns.
Labor cost Z occurrence !f½detection time The availability of a repairable component is approxi-
!labor rate !no: of operators mated as expressed in Eq. (5), if after each repair ‘as good as
new’ is assumed [12].
C ½fixing time !labore rate
† Availability. Average probability that an item will
!no: of operators C ½delay time perform its required function under given conditions at
time.
!labor rate !no: of operatorsg (2)
MTTF MTTF
Availability ðAÞ Z Z (5)
Component replacement due to failure is considered as MTTF C MTTR MTBF
material cost. Material cost is obtained using the following
equation: † MTBF (Mean Time Between Failures). MTBF is a basic
Material cost Z occurrence !cost of part (3) measure of reliability for repairable items. It can be
described as the number of hours that pass before a
Opportunity cost is the cost that incurs when a failure component, assembly, or system fails. It is a commonly
inhibits the main function of the system and prevents any used variable in reliability and maintainability analyses.
creation of value. Opportunity cost is the cost incurred when † MTTR (Mean Time to Repair). MTTR is the average time
a failure inhibits the main function of the system and required to perform corrective maintenance on all of the
prevents any creation of value. Opportunity cost is removable items in a product or system. This kind of
calculated using the following equation: maintainability prediction analyzes how long repairs and
Opportunity cost maintenance tasks will take in the event of a system failure.
† MTTF (Mean Time To Failure). MTTF is a basic measure
Z down time !hourly opportunity cost (4) of reliability for non-repairable systems. It is the mean time
expected until the first failure of a piece of equipment.
where MTTF is a statistical value and is meant to be the mean over
Down time Z fdetection time C fixing time C delay timeg a long period of time and large number of units.

Failure rate, m, can be expressed in the following way:


2.3. Life cost-based with Monte Carlo simulation 1
mZ (6)
MTBF
Life Cost-Based FMEA uses point estimation for its
analysis. The danger with using point estimation is the The method mentioned above can be applied to
potential for misinterpretation of the average numbers. predicting the availability for systems or subsystems.
Plans based on average conditions are incorrect since one Availability of a subsystem that has many components of
does not know if the condition has reached the upper or the same type in series can be modeled using the
lower thresholds. A sensitivity analysis on the estimates will following equation:
provide better confidence in the result and make for a better AvailabilitySubsystem
understanding of which variables are the cost drivers.
A Monte Carlo simulation is applied to the Life Cost- Z ðAvailability1Component Þno: of components (7)
Based FMEA to perform a sensitivity analysis on the
variables associated to failure cost: occurrence, detection We can predict failure frequency for a given time once
time, fixing time, and delay time. A triangular distribution the availability has been calculated. Downtime for the
using minimum, mode, and maximum value was used. system is calculated using the following equation:
There are many distribution systems one can use for Downtime Z ð1 K availabilityÞ !operation time (8)
the simulation; however, with limited past history data
and using estimated variables, a triangular distribution was Knowing MTTR from empirical date, the average failure
selected. frequency for a given time can be predicted using
S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188 183

the following equation: Table 2


Downtime of accelerator due to magnet failure
Downtime
Failure Frequency Z (9) Events Total Min Max Avg
MTTR downtime
Although empirical data for average frequency, detection Solid 6 25.8 1.8 11 4.3
time, fixing time and delay time can be obtained it is still wire
dangerous to use point estimation. Any decisions based on Water 70 699.2 0.1 32 10.0
cooled
average conditions could be incorrect since one does not
Total 76 725 9.5
know if the condition has reached the upper or lower
thresholds. A sensitivity analysis on the estimates will Units: hour.
provide better confidence in the result. A Monte Carlo
simulation is applied to the Life Cost-Based FMEA to the NLC are categorized into two fundamental designs:
perform a sensitivity analysis on variables related to failure water-cooled and solid wire electro magnets.
cost: occurrence, detection time, fixing time, delay time
and material cost. A triangular distribution using minimum, 4.1. Electromagnet
mode, and maximum value was used. The results are
discussed in Section 5 of this paper. History of magnet failures from 1997 to 2001 was
collected using the SLAC CATER system. Table 2 shows a
total of 76 incidents where a beam line had to be shutdown
due to electromagnet failures. Ninety percent of the failure
4. Actual application incidents and 96% of the total downtime were associated
with water-cooled electromagnets.
SLAC is a national research laboratory that is charged Table 3 shows the five most common failure modes with
with investigating the most basic elements of matter. electromagnets and its frequency and downtime. Insulation
Engineers at SLAC and other labs are currently designing failures are related to magnets becoming shortened. The
NLC that will be 20 miles long, 10 times longer than the insulation material around the coil becomes degraded due to
current linear accelerator at SLAC. The proposed NLC has a radiation or thermal effects. Water leaks from the cooling
proposed 85% overall availability goal, the availability system is a major problem with electromagnets. Water leaks
specifications for all its 7200 magnets and their 6167 power are mainly due to failures in flexible hoses and copper
supplies are 97.5% each. SLAC intends to operate the NLC corrosion. Human errors range from not following pro-
24 h per day, 7 days a week for 9 months a year. Thus, all of cedures or forgetting tasks. Connector failures are mainly
the electromagnets and their power supplies must be highly mechanical failures due to mechanical and thermal cycling.
reliable or quickly repairable to minimize interruption of the If the NLC were to be built using all electromagnets, the
particle physics research program. NLC would require 2202 solid wire and 4965 water-cooled
SLAC keeps a history of all failures for the past 15 years electromagnets. All 7167 electromagnets are needed to be
on an online database called the Computer Aided Trouble available for the accelerator to run. Thus, if one magnet
Entry and Reporting (CATER) system [13]. MTBF for fails, the whole accelerator will come to a stop. We can
failures modes identified in FMEA is acquired using data estimate the availability of the magnet system (ASys) using
from SLAC’s failure database and design experience. the following equation
Occurrence or probability for certain failure modes can be
determined through MTBF. NLC requires 7167 magnets to ASys Z ðA1C Þn (10)
control its particle beams. Two different technologies could where A1C is the availability of one component and n is the
be used for the magnets: electromagnet or permanent total number of components in the system.
magnets. Table 4 shows the different types of beamlines and their
An electromagnet’s strength is varied by changing the durations during the period from 1997 to 2002. The second
electric current in the coils. Thus, a power supply is required column indicates the type of line running during the period.
as part of the system. SLAC has been using more than 3000
electromagnets over the past 30 years and their failure data Table 3
Failure frequency and downtime of electromagnets
are readily available. Another competing technology uses
permanent magnets without any current. Permanent mag- Failure mode Events Min (h) Max (h) Avg (h)
nets are simpler in design and the initial manufacturing cost Insulation 29 0.2 27.2 8.82
maybe smaller. However, the technology risk has not been Water leak 22 1 32 9.7
ascertained for adjustable permanent magnets. Water blockage 5 0.5 7.5 3.92
This paper describes the utilization of empirical data to Human error 5 0.7 6 2.5
Connector 3 1 3.2 1.733
estimate failure/risk cost for electromagnets. One hundred
Other 12 0.9 10.2 5.8
and twenty-nine different types of electromagnets for
184 S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188

Table 4
Run time of water cooled electromagnet

By run time (water cooled magnets)


Date Line Run Magnets Magnet # Failures MTBF TR MTTR Availability
hour hours 1 Mag
2/4/97–4/30/97 Linac/BSY 1547 520 804,440 1 804,440.0 0.2 0.20 0.999999751
5/1/97–6/8/98 SLC 8828 2104 18,574,112 32 580,441.0 469.5 14.67 0.999974724
7/10/98–7/31/98 HER and LER 575 2433 1,398,975 2 699,487.5 9 4.50 0.999993567
10/30/98–12/1598 HER and LER 1040 2433 2,530,320 6 421,720.0 40.1 6.68 0.999984152
1/15/99–2/22/99 HER and LER 844 2433 2,053,452 4 513,363.0 15.6 3.90 0.999992403
2/24/99–5/1/99 Linac 1461 520 759,720 2 379,860.0 26.1 13.05 0.999965646
5/1/99–11/29/99 HER and LER 4797 2433 11,671,101 7 1,667,300.1 65.65 9.38 0.999994375
1/12/00–10/31/00 HER and LER 6624 2433 16,116,192 7 2,302,313.1 34.6 4.94 0.999997853
1/10/01–12/31/01 HER and LER 7411 2433 18,030,963 7 257,5851.9 37.9 5.41 0.999997898
Sum 75,475,138 70 701.70
Average 1,078,216 10.02 0.9999907029

Magnets System

Actual PEP II 2433 Availability 0.984009931 0.999993375


Actual SLC 2104 Availability 0.948207061 0.999974724
Predicted NLC 4965 Availability 0.95488886

Forecast for NLC

Operation (h/yr) 6480


Expected downtime 292.3 h/yr
Occurrence/yr 29.2

The third column shows the number of water-cooled the following equation:
electromagnets for that particular line. The fifth column is AMSys Z ASM !AWM Z 0:9987 !0:9549 Z 0:9536 (11)
the product of run hour and the number of magnets: magnet
hours. The sixth column indicates the number of failures
identified during that particular period. The MTBF in the ASM availability of solid wire magnet
seventh column is a result of magnet hours divided by the AWM availability of water-cooled magnet
number of failures. The eighth column indicates the total
repair time for those failures in that period and the ninth Thus, this would fall short of the 97.5% availability goal
column is MTTR. Based on these numbers, the availability if the design of the new magnets does not eliminate the
of any one magnet in a beamline can be calculated. root cause of the observed failures. A summary of the
The average availability of one water-cooled magnet at availability is shown in Table 5.
SLAC is found to be 0.9999907. This example predicts the overall failure for the NLC,
but one can predict failures for particular types of failure
The availability of the NLC’s electromagnet subsystem
(insulation, water leak, water blockage, mechanical, or
can be estimated using Eq. (7). Assuming the reliability of
human error) using the same methodology.
each individual magnet is 0.9999907 the availability for
4965 water-cooled electromagnets would be 0.9548. 4.2. Power supply
However, this is lower than the target value of 97.5% for
the magnet subsystem. Therefore, the magnet designers The power supplies that provide the electric current to
know they must improve the reliability of the magnets they the electromagnets can be categorized into two main
design for NLC over the SLAC magnets. Given 6489 h of
operation time per year, the expected downtime of the NLC Table 5
due to electromagnet failure is 292 h/yr. Since the average Predicted availability of electromagnets for NLC
MTTR is 10 h, we can estimate the number of failures for a Type Solid wire Water-cooled
given year to be 29 occurrences. No. of Magnets 2202 4965
Availability of solid wire magnets can be calculated in Availability 0.9987 0.9548
the same manner. The expected number of failures for solid Expected downtime 8.3 292
(h/yr)
wire magnets in the NLC is twice a year. The overall
Occurrence per year 1.9 29.2
availability of the NLC magnet system is obtained using
S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188 185

Table 6
Downtime of accelerator due to power supply failure

Type of Events Total down Max Min Avg down


PS time time
Large 92 178 11 0.1 1.93
Small 70 88.7 11.5 0.2 1.27
Total 162 266.7 32 0.1 1.65

Units: hour.
Fig. 2. Electromagnet system.

categories: small (!12 A, 50 V) and large (O12 A, 50 V).


A summary of the SLAC power supply failures from the Fig. 2 schematically. The larger water-cooled magnets will
CATER system between 1997 and 2001 is shown in Table 6. have redundant power supplies since the availability of a
The total number of failures is 2.5 times greater than the single power supply is too low. The accelerator will
number of electromagnet failures but the total downtime is shutdown if any one of the 7167 magnets or 6167 sets of
less than half of the electromagnet failures. This is because power supply fail. Thus, the availability of the system is
the average downtime for power supply is only 1.65 h as the product of the magnet and power supply availabilities.
opposed to 9.5 h for the electromagnet. Only failures that
required the accelerator to shutdown were considered. ASys Z AMSys !APSSys Z 0:9536 !0:986 Z 0:94
Availability, MTBF, and occurrence for the power
supplies can be estimated following the same steps as in Expected downtime
the electromagnets. The results are shown in Table 7.
Z ð1 K availabilityÞ !operation hour=yr
The overall availability of the power supply system
(APSSys) is the product of the two types of power supplies: Z 0:0597 !6480 h=yr Z 387 h=yr
small and large.
APSSys Z ASPS !ALPS Z 0:988 !0:938 Z 0:927 Occurence Z expected downtime=MTTR

Z 387ðh=yrÞ=9:6 ðhÞ Z 40:3=yr


ASPS availability of small power supply
ALPS availability of large power supply Having looked into the two major sub systems for the NLC,
magnet subsystem and the power supply subsystem, the life
This is far short of the 97.5% availability requirement. Thus, cycle cost of the whole magnet system can be analyzed.
the reliability of the power supply has to be increased. One way
of achieving this is to design redundancy in the system. Since 4.4. Life Cost-Based FMEA
the small power supply has a high availability rate, we will
consider having redundancy only in larger power supplies to A Life Cost-Based FMEA sheet, as shown in Table 1,
minimize cost. With redundancy in large power supplies was completed for the electromagnet. First, the origin of the
(ALPS), the power supply availability becomes 0.986. failure and the detection stages were identified for each
scenario. Failure frequency was assigned with respect to the
Improved APSSys Z 0:988 !0:9975 Z 0:986
availability model discussed in Section 3. Experts in their
The expected downtime due to power supply failure is respected fields gave design, manufacturing, and installation
6480 h!(1K0.9855)Z93.9 h/yr. Using the average fixing failure frequencies.
time for the power supply, 1.5 h, and the average number of Labor cost is obtained using Eq. (2) with $60/h for labor
failure during the year is 93.9/1.5Z62 events/yr. rate and assumes two operators are required to detect and fix
the problems. Repairs for an electromagnet failure can cost
4.3. Electromagnet system from a simple replacement of water hose to replacing the
whole magnet. Material cost for a simple replacement of
The electromagnet system has power supplies that hoses can range from $35 to $70. Replacing the whole solid
control the electric field for each magnet as shown in wire magnets can range in cost from $400 to $2000 and water-
cooled magnets can range from $4000 to $30,000 depending
Table 7 on the size of the magnet. Power supply repairs usually only
Predicted availability of power supplies for NLC require the electronic boards to be switched and parts cost for
Size Small Large boards range from $300 to $700 depending on the size.
No. of power supplies 2785 3382 SLAC estimates the lost opportunity due to shutdown to
Availability 0.988 0.938 be anywhere between $10,000 and $50,000 per hour for the
Expected downtime (h/yr) 77 400
MTBF (h) 105 32
NLC. $10K is estimated if only direct labor costs are
considered, $25K when direct labor and wasted energy costs
186 S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188

Fig. 3. Monte Carlo simulation of labor and material cost for electromagnet.

Table 8
Predicted life cycle failure cost of electromagnets for the NLC for 30 years

Correctors Water cooled Electro magnet


Probability Probability Probability
5% 50% 95% 5% 50% 95% 5% 50% 95%
Labor cost $0.300 $0.395 $0.532 $1.200 $1.490 $1.760 $1.500 $1.885 $2.292
Material cost $0.065 $0.082 $0.102 $0.900 $1.120 $1.360 $0.965 $1.202 $1.462
Sub total $0.37 $0.48 $0.63 $2.10 $2.61 $3.12 $2.465 $3.087 $3.754
Opportunity cost $10K $6.0 $7.8 $9.7 $72 $90 $109 $78.0 $97.8 $118.7
$25K $15.0 $19.7 $24.2 $175 $225 $272 $190.0 $244.7 $296.2
$50K $30.0 $38.0 $48.0 $350 $450 $543 $380.0 $488.0 $591.0

Units: million.

are considered, and $50K when the cost of building the NLC large power supplies is still quite high because the power
is amortized over a 30-year period in addition to the labor supply electric boards have to be replaced regardless of the
and energy cost. Thus, the overall opportunity cost was shutdown of the accelerator.
calculated for all three values. The magnet system requires the electromagnets and
A Monte Carlo simulation is applied to the Life Cost- power subsystem to both be working. Thus, the life cycle
Based FMEA to consider the sensitivity of variables failure cost of the subsystem is the sum of electromagnet
associated to failure cost: frequency, detection time, fixing and power supply failure cost as shown in Table 10. The
time, delay time, and parts cost. Fig. 3 shows result of the actual labor and material cost is a small fraction of what the
simulation for labor and material costs for the different total opportunity cost might be, even using the lowest
confidence levels. A 30 year predicted failure cost for the opportunity cost per hour, $10K/h.
electromagnet is summarized in Table 8. As shown in the
table, opportunity cost can be 30–150 times greater than
the labor and material cost. 5. Discussion
The estimated failure cost for the system of power
supplies is summarized in Table 9. As predicted in Table 7, As derived in Section 4, availability for the electromag-
the availability of large power supplies is pretty low, 0.938. net system falls short of the target goal of 97.5%. To
Thus, redundancy is assumed for the large power supplies to increase the availability of the water-cooled magnets for
meet the availability goal. Material and labor failure cost for the NLC, two measures can be taken: reduce MTTR or

Table 9
Life cycle failure cost of power supply of 30 years Table 10
Life cycle failure cost of electromagnet system
Small Large Total
Failure cost
Labor cost $0.39 $1.90 $2.29
Material cost $0.92 $7.20 $8.12 Labor cost $4.2M
Sub total $1.31 $9.10 $10.41 Material cost $9.3M
Opportunity cost $10K $23 $6 $29 Sub total $13.5M
$25K $59 $15 $74 Opportunity cost $10K $126.8M
$50K $117 $30 $147 $25K $318.2M
$50K $635M
Units: million.
S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188 187

A4965M Z ð0:9999953Þ4965 Z 0:977

AMSys Z AWM !ASM Z 0:977 !0:999 Z 0:976

A4965M availability of 4965 magnets


AMSys availability of the magnet system

A summary of the Monte Carlo simulation of the Life


Cost-Based FMEA with improved MTTR of 5 h for failures
that occur during operation period is shown in Table 11. A
$600,000 dollar savings can be expected over 30 years in
labor and material costs if the fixing time is reduced by 50%.
However, the incentive to reduce fixing time is more evident
in opportunity cost. Comparing Tables 8 and 11, a 40%
reduction in opportunity cost can be expected through
reducing the time spent on fixing electromagnets.
Fig. 4. Top failure costs.
Another approach to enhancing availability is to use
reliability allocation and optimization method that enables
increase the reliability of the electromagnets. The average designers to determine which components to increase its
MTTR for water-cooled electromagnet is 10 and 2 h for the reliability to meet the overall reliability goal of the entire
solid wire magnets as determined from empirical data at system [15]. Reliability allocation is a detailed analysis
SLAC. Referring to Table 3, the average fixing time for looking at each individual component to determine the
insulation and water leak is 8–10 h. reliability enhancement for each component at a detailed
Fig. 4 shows the top four failure costs with respect to design level. Since the authors are investigating risk at the
its root cause. Opportunity cost of $25,000 per hour was early design concept stage, we did not feel the need to apply
used for this analysis. Water leak has the highest failure reliability allocation to the methodology.
cost with a total cost of over one hundred million dollars
with a 50% probability. Thermal, radiation, and mechani-
cal are the following high cost failures. Thus, we will
investigate and suggest recommendations regarding water 6. Conclusions
leak failures.
The time to fix water leak failures can range anywhere This paper demonstrated the systematic use of empirical
from as little as 1–32 h. A bad water leak can take up to 32 h data in performing Life Cost-Based FMEA and how it can
to dry out the magnets. Design improvements to shorten the improve the reliability, maintainability, and life cycle cost
fixing time of replacing the fittings and coils will of complex systems such as a linear particle collider. Life
significantly decrease MTTF for water leaks. An average Cost-Based FMEA aids not only design improvements and
50% reduction in MTTR, from 10 to 5 h, for water-cooled concept selection, but it also allows one to improve and plan
electromagnets will increase the availability of magnets to preventive and scheduled maintenance of components.
97.6% as shown in the following equations: Thus, Life Cost-Based FMEA has three main benefits:
(1) Estimation of life-cycle cost, (2) FMEA, and
(3) Service Mode Analysis (SMA) [16].
MTTF 1; 078; 216 The proposed method inherently captures a system’s
A1M Z Z Z 0:9999953
MTTF C MTTR 1; 078; 221 life-cycle costs related to component failures during design,
Table 11
Predicted life cycle failure cost of electromagnets with 50% reduction in fixing time

Correctors Water Cooled Electro Magnet


Probability Probability Probability
5% 50% 95% 5% 50% 95% 5% 50% 95%
Labor cost $0.300 $0.395 $0.532 $0.600 $0.880 $1.150 $0.900 $1.275 $1.682
Material cost $0.065 $0.082 $0.102 $0.900 $1.120 $1.360 $0.965 $1.202 $1.462
Sub total $0.37 $0.48 $0.63 $1.50 $2.00 $2.51 $1.865 $2.477 $3.144
Opportunity cost $10K $6.0 $7.8 $9.7 $43 $52 $67 $49.1 $59.8 $76.2
$25K $15.0 $19.7 $24.2 $103 $130 $160 $118.0 $149.7 $184.2
$50K $30.0 $38.0 $48.0 $206 $260 $320 $236.0 $298.0 $368.0

Units: million.
188 S.J. Rhee, K. Ishii / Advanced Engineering Informatics 17 (2003) 179–188

manufacturing, installation, and operation. Designers can Cherrill Spencer and John Cornuelle for their valuable time
readily incorporate the changes in the model to estimate an spent on this research, and other SLAC staff members who
improved life cycle cost. The root causes directs designers provided us with useful technical information.
to focus their efforts on problem systems, components, and
processes.
Complex systems usually have set target availability.
One means to achieve the target is to increase all subsystems References
reliabilities. However, guaranteeing higher reliability often
incurs cost increases. Another solution is to schedule [1] Stamatis DH. Failure mode and effect analysis. Milwaukee, WI: ASQ
preventive maintenance. Our proposed methodology maps Quality Press; 1995.
[2] Lee B. Using Bayes belief networks in industrial FMEA modeling and
allow comparisons of different availability enhancement
analysis. Proceedings of International Symposium on Product Quality
measures and trace analysis in terms of cost, a widely and Integrity, Philadelphia, PA; 2000.
accepted measure of risk. [3] Kmenta S, Ishii K, Scenario-based FMEA: a life cycle cost
The authors agree that extracting relevant knowledge perspective. Proceedings of ASME Design Engineering Technical
from pre-existing data (CATER system) is hard work Conference, Baltimore, MD; 2000.
[4] Palady P. Failure modes and effects analysis; predicting and
because it collects data without the purpose of improving
preventing problems before they occur. Florida: PT Publications;
reliability or serviceability. It is evident that reliability and 1995.
serviceability should be considered when maintenance [5] He D, Adamyan A. An impact analysis methodology for design of
management systems are put into place. Many times the products and processes for reliability and quality. Proceedings of
management system records data in a way that it cannot be ASME Design Engineering Technical Conference, Pittsburgh, PA;
2001.
used effectively. One could apply data mining techniques to
[6] Tarum CD. FMERA—failure modes, effects, and (financial) risk
improve feedback from experience to extract relevant data analysis. SAE World Congress, Detroit, MI; 2001.
from tons of data that have been recorded [17]. The example [7] Gilchrist W. Modeling failure modes and effects analysis. Int J Quality
presented in this paper used a semi-manual sorting Reliab Manage 1992;10:16–23.
technique to extract the relevant data. [8] Ben-Daya M, Raouf A. A revised failure mode and effects analysis
Life Cost-Based FMEA can also provide a fair model. Int J Quality Reliab Manage 1996;13:43–7.
[9] Bevilacqua M, Braglia M, Gabbrielli R. Monte Carlo simulation
comparison between competing designs of subsystems. approach for modified FMECA in a power plant. Quality Reliab Eng
The case study presented in this paper considered only Int 2000;16:313–24.
the currently used magnet technology. The proposed [10] MIL-STD-1629A, BS 5760 Part 5.
methodology may not simply extrapolate to new and/or [11] SAE ARP-4293: Life cycle cost—techniques and applications.
unproven technology, because empirical or expert knowl- [12] Birolini A. Quality and reliability of technical systems, 2nd ed. Berlin:
Springer; 1997.
edge may not be available. Thus, future research lies in [13] Sass R, Shoaee H. CATER: an online problem tracking facility for
estimating uncertainty variables (e.g. frequency, detection SLC. Proceedings of Particle Accelerator Conference, Washington
time, fixing time, and delay time) using available com- DC, 1993.
ponent data, and extrapolating them to higher subsystem [14] Rhee S, Ishii K. Life cost based FMEA incorporating data uncertainty.
levels. Hybrid use of empirical and analytical data will Proceedings of ASME Design Engineering Technical Conference,
Montreal, Canada; 2002.
present significant new challenges.
[15] Mettas A. Reliability allocation and optimization for complex
systems. Proceedings of IEEE Reliability and Maintainability
Symposium, Philadelphia, PA; 2001.
Acknowledgements [16] Gershenson J, Ishii K. Design for serviceability. In: Kusiak A, editor.
Concurrent engineering: theory and practice. New York: Wiley; 1992.
p. 19–39.
This research has been supported by the Department of
[17] Manago M, Auriol E. Using data mining to improve feedback from
Energy contract, DE-AC03-76SF00515. The authors would experience for equipment in the manufacturing and transport
like to thank the Stanford Linear Accelerator Center for industries.: Institute for Operations Research and the Management
providing the opportunity for this research, and especially Scienece; 1996.

You might also like