Finding
Value
How to Determine the Value of
Reliability Engineering Activities
Fred Schenkelberg
Finding
Value
How to Determine the Value of
Reliability Engineering Activities
Fred Schenkelberg
Los Gatos, California
2014
Copyright © 2014 Fred Schenkelberg
Licensed under the Creative Commons
Attribution-NonCommercial-NoDerivatives
4.0 International License.
[Link]
Feel free to email, tweet, blog, and pass this ebook
around the web but please don’t alter any of its
contents when you do. Thanks!
If you find this work of value to you, consider
purchasing a copy and supporting the work.
If you have purchased a copy, Thank you!
FMS Reliability Publishing
15466 Los Gatos Blvd #109-371
Los Gatos, California 95032
[Link]/publishing/
Printed in the United States of America
ebook ISBN: 978-1-938122-02-6
paperback ISBN: 978-1-938122-03-3
Contents
Introduction1
Reliability Value 3
FMEA: How to Find Value 11
Methods to Estimate Accelerated Life Testing Value 21
An Estimate of Highly Accelerated Life
Testing Value 29
Estimating the Value of Derating 39
Reliability Prediction Value 45
Increasing Value 59
11 Ways to Find Reliability Value 63
Notes75
Introduction
Introduction
An obvious result of good reliability engineering is the lack of field
failures. Connecting your work to the results is not always obvious.
In todays lean organizations everyone has to provide tangible value.
Yet, if the product is doing well, how do you show your ongoing
contribution to the organization?
Reliability engineering may increase the cost of a product or
recommend expensive product testing. Justifying these expenses is
often based on the chance of improved product reliability.
It is the quantification of value due to specific reliability engineering
actions that enables you to articulate your worth to an organization.
This short book explores how to calculate the value of reliability
engineering activities. We explore ways to estimate value for use in
engineering proposals.
We know that a reliable product provides value to the customer, it also
is a value to you and your organization.
Here you will learn how to connect specific reliability engineering work
to the real value created.
1
finding value – how to determine the value of reliability engineering activities
2
Reliability Value
Reliability Value
What is reliability management?
What is reliability engineering?
Would a product design or an organization benefit by focusing on
reliability management and engineering?
What is the value of a focus on reliability?
Any organization that develops and produces products has limited
resource. These may be talent, capabilities, time, funding, or some
combination of these.
Yet, the goal to create a product that meets customer expectations
includes the concept of product reliability.
The product should provide the expected functions over time, without
failure.
This expected product reliability exists, even if the design requirements
and advertising do not explicitly mention product reliability.
3
finding value – how to determine the value of reliability engineering activities
For example, consider a laptop that needs a new power supply.
When this situation came up, my first thought was to consider how old
the machine was. Was it still under warranty?
Then my thoughts turned to the inconvenience of either being without
my laptop during the repair period or the hassle of moving over to a
new machine.
If the machine was only a few months old, it would likely still be under
warranty—yet my dissatisfaction would be higher.
It shouldn’t have failed so soon.
If the machine was five years old, that would be a different story.
I’d have had many years of use and, if this was the first failure, I’d have
gotten a lot of value.
Besides, it may well be time to upgrade to a new machine. The
inconvenience of a repair or purchase of a new machine, although not
totally alleviated, is still much less.
4
Reliability Value
Value of Product Reliability
The primary value of product reliability is in meeting the customer’s
expectation that the product will work as intended for sufficient time.
The market rejects products that fail often and desires products that
‘just work.’
Creating a reputation for a reliable product assists in increasing sales.
An extension of the value that consumers place on reliability is
their willingness to pay a premium for products with high reliability.
Automobiles, computers, printers, appliances, and test equipment
are all areas where products of known high reliability can command a
premium.
Paying a premium is worth it: The cost of downtime during a failure
more than outweighs the additional purchase expense.
For the business creating a reliable product, it creates value in a
similar manner.
Products that are sought after and command a price premium lead to
higher sales and higher profit margins.
5
finding value – how to determine the value of reliability engineering activities
Additionally, the lower failure rates reduce warranty expenses, which
further increases the business’s profit margin.
Yes, it may cost more in materials to create a durable product, but it
returns rewards of higher customer satisfaction, market share, and
profit margin.
Reliability Engineering
Reliability is
the probability that an item will perform a required function
without failure under stated conditions for a stated period of
time.1
Reliability engineering is an engineering field that deals with the study,
evaluation, and life-cycle management of reliability.
Reliability engineering includes the use of statistics, data analysis,
experimental design, customer and environmental surveys, component
and product testing, failure analysis, design, manufacturing,
procurement, and, at times, marketing and finance.
6
Reliability Value
Reliability engineers must have a broad set of skills, and the proper
application of reliability tools and techniques generally permits an
organization to create a reliable product.
The role of a reliability engineer spans a variety of tasks and
disciplines.
While some reliability engineers will specialize in one area of the field,
say accelerated testing, others may find a role that involves nearly
every function within an organization.
The ability to influence and create a product that meets the customer’s
reliability performance expectations is both challenging and
rewarding.
Reliability Management
Oversight and control of reliability activities is a key management role.
Some organizations have a dedicated reliability manager, others
identify a senior reliability engineer, and in others reliability
management is part of the organizations management functions.
7
finding value – how to determine the value of reliability engineering activities
There is no one right way to organize to accomplish improved product
reliability.
It is more important to focus—across the organization—on the impact
of decisions on the resulting product’s reliability performance.
The management of reliability, like reliability engineering, may involve
working closely with many functions throughout an organization.
Reliability engineering and management are very similar.
The former works to implement activities and analysis that enable the
creation of a reliable product. The latter does the same though the
allocation of resources to enable the right activities and analysis.
The organization that includes reliability considerations (i.e.,
requirements, predictions, risks, evaluations, and analysis)
deliberately and uses the information to guide decisions across the
organization will create reliable products.
Those that ignore or isolate reliability to a limited role within the
organization are less likely to create a reliable product.
8
Reliability Value
The actual individual titles are less important than the specific
reliability engineering activities and decisions.
Reliability engineering skills are part of any engineering discipline;
with some practice and encouragement nearly all engineers have the
capability of learning the needed skills.
The necessary management skills are similar to any other product-
producing organizational set of skills.
The ability to coordinate activities, allocate resources, and focus
on reliability is augmented by a solid understanding of reliability
engineering tools and techniques, just as with any other management
task.
9
finding value – how to determine the value of reliability engineering activities
10
FMEA: How to Find Value
FMEA: How to Find Value
Failure modes and effect analysis (FMEA) is a tool that works to
prevent process and product problems before they occur.
One way to define FMEA is as an organized brainstorm. In the process
the FMEA team examines a product or process and asks, `What could
go wrong?’ Then the team systematically determines and ranks each
failure mode with:
• the severity of the problem when it occurs,
• the probability of the problem occurring, and
• the ability to detect the problem before it occurs.
Good design engineers think about how the design could fail and
improve the design. FMEA provides a structured team approach to
further improve the design.
Cost of the FMEA Study
The cost of an FMEA is generally easy to determine. It is the cost of the
talent on the team and any specific experiments or expenses related to
the FMEA study.
11
finding value – how to determine the value of reliability engineering activities
The cost of changing a design based on the study is not included,
provided the design team would already be working to improve the
product. The FMEA results provide direction for improvement work,
thus focusing the effort, rather than adding to the effort.
Let’s say the team consists of five people and they spend a day working
on the FMEA study. Further, let’s say each engineer has a loaded cost
to the organization of $1,000/day. Thus the engineering time cost of the
study is $5,000.
If the team does a few experiments or has specific expenses, that
would add to the cost. In this simple example, let’s say they used one
prototype in a short experiment to answer a few outstanding questions.
To make it simple, assume the additional cost is $5,000.
The total cost (talent and expenses) would thus be $10,000.
Additional cost may include training and facilitation support.
Sources of FMEA Value
So, what do we get for the investment of time and resources in a simple
design FMEA?
12
FMEA: How to Find Value
The return on investment for an FMEA study largely depends on when
it is done and how well the team implements the action items (the
outcomes of the study).
It also depends on existing knowledge about product or process
problems. If we already have a long list of issues and an easy way to
prioritize them, say by analyzing field returns for an existing product,
then FMEA may not add any new information.
If, however, the product is using a new design approach or the process
is manipulating a new material, then we may not have a well-crafted
list of issues.
FMEA is a great tool when there are unknowns or uncertainness. It is
also commonly used by teams to organize and manage a large number
of risks.
It is not appropriate to expect any value when the study is always done
despite existing knowledge. If the study only lists what we already know,
it has little value.
When the tool is applied in a situation early in the design process of a
new product or process, with new design approaches, architecture,
13
finding value – how to determine the value of reliability engineering activities
or materials, and the risks of failure are either unknown or uncertain,
then FMEA may provide significant value.
In my experience, I’ve seen FMEAs create value in three basic ways:
1. by identifying and removing faults,
2. by prioritizing design improvement work, and
3. by providing cross-department coordination.
Let’s explore each aspect briefly.
Identifying and Removing Faults
This is an obvious value-added function. Let’s say the team already
knew of about ten problems that needed addressing in the design. Then
the FMEA added an additional issue to resolve.
An estimate of the impact of that issue if unresolved is part of the
prioritizing process in an FMEA and thus provides a rough estimate
of the percentage of products that would have failed in the hands of
customers.
14
FMEA: How to Find Value
The number of failures times the cost per failure provides the avoided
cost.
For a simple example, let’s assume we plan on selling 10,000 units of
a particular design. Each failure costs our organization $500. In the
FMEA we estimated that the problem would cause about 1% of the
units to fail.
Now 1% of 10,000 is 100 failures. At a cost of $500 each, that would be a
cost of $50,000.
Normally, a well-done FMEA will reveal many previously unknown
problems. The issues may impact most or very few products, yet the
same basic logic applies. Avoiding field failures avoids the cost of
failures occurring.
In this example, we have considered only direct costs, such as warranty
replacement or repair costs. In reality, the failure cost may include
degradation of brand image or customer satisfaction.
15
finding value – how to determine the value of reliability engineering activities
Prioritizing Design Improvement Work
Even simple designs may have many hundreds of potential failure
modes.
Solving the problems takes time and resources and solving the wrong
problems adds little value. Focusing on the most severe and most likely
to occur issues tends to provide the most value.
I’ve worked with hundreds of design engineers and have asked many of
them about what could go wrong with their current design. Each easily
can provide several issues that should be addressed before finalizing
the design.
One problem that FMEA helps to solve has to do with limiting the list of
priorities. A design team of five people all working on the same product
can easily create a list of top priorities that most likely will not overlap.
The design team generally cannot have five top priorities.
The larger the team and the more complex the product or process,
the longer is the list of problems to solve. Without some way to
organize priorities, the team easily could be working aimlessly to solve
individually determined top priorities.
16
FMEA: How to Find Value
This generally is not an efficient use of design talent.
FMEA is not the only way to provide order, yet it does provide a well-
thought-out, reasonable way to focus the design team to solve the top-
priority problems.
Most designers quickly consider the problems that will effect all
products or have dangerous failure modes. The balance and priorities
become more difficult when the chance and severity are not as
significant.
The FMEA study enables the team to articulate the priorities clearly.
The value is found by shifting focus to the most important issues.
Let’s say a team identifies 100 issues that could improve the design.
The program has the resources to solve 20 issues before product
launch.
Which 20 should the team focus on?
Using the basic concept of 80% of field issues are caused by 20% of the
design flaws, if we select the right 20 issues we can avoid the majority
of field issues.
17
finding value – how to determine the value of reliability engineering activities
If we select randomly, we may only address 50% of the field issues.
Given simple engineering judgment this will most likely not identify and
solve the major issues.
So let’s say we’re selecting the issues at the margin of what we can
address and solve, that the last four issues identified to solve in the
FMEA study would have been randomly selected otherwise. The
resulting benefit is the avoided field failures depend on the difference
in expected impact to the cost of failure, yet even this small change in
selection of issues to address provides enough return on the FMEA
investment to provide value to the team and customer.
Of course, the prioritizing creates a list with diminishing return on the
effort to solve it prior to shipping. Yet, the net effect adds value.
Cross-Department Coordination
In most product or process design work, the complexity of the system
requires teams focusing on mechanical, electrical, and software
elements of the design.
One of the dangers that each team faces is scarce resources. If you’ve
considered prototype allocation you can appreciate what this means.
18
FMEA: How to Find Value
Each team may fully understand the risks that may cause field failures,
yet they often lack the understanding from the other teams. Each team
may view the design challenges they face as the most important.
One benefit that is difficult to quantify is the ability of the study to
establish cross-team discussions about priorities for the available
resources.
Reducing the cross-team competition frees engineering and
management resources to focus on and solve the most important
problems. It makes for a better work environment, which is a real value
in itself.
Although difficult to quantify, this improved work environment has a
tangible benefit.
Summary
FMEA is a powerful analysis tool.
When used wisely projects that benefit from the systematic
identification and prioritizing of potential failure modes, FMEA will
provide significant value.
19
finding value – how to determine the value of reliability engineering activities
20
Methods to Estimate Accelerated Life Testing Value
Methods to Estimate Accelerated Life Testing
Value
Here is an example of how to determine the future value of a specific
reliability task.
Many of us face the challenge of how to justify spending product
development resources to provide insights and information to the rest
of the team.
Accelerated life testing (ALT) is particularly difficult, as it is time
consuming, expensive, and at times statistically complex. Having
a clear method to estimate the value serves your career and
the organization well, as both will benefit from making the right
investments.
ALT and Market Share
A design team working on a medical device understands that market
share is related to product reliability.
21
finding value – how to determine the value of reliability engineering activities
The current product performs adequately, yet has the highest field
failure rate of similar products. Customers complain about the poor
reliability, and the market share reflects its comparative reliability
ranking.
The most reliable product is also the one with the highest market
share.
The design challenge is to create a product that is more reliable than
that of the competition, at about the same price point, and if possible
with improved functionality.
The early concepts all include a novel design using an unproven
sealing material (in terms of reliability). The uncertainty suggests the
implementation of an accelerated life test to estimate the expected
product reliability.
Achieving a higher reliability is expected to result in more than tripling
the market share in the first year. This would result in sales of the
$3,000/unit product to jump from 10,000/year to approximately
30,000/year, meaning an additional $60 million in revenue.
Furthermore, the increase in sales would require more than doubling
the manufacturing capacity, at a cost of $5 million.
22
Methods to Estimate Accelerated Life Testing Value
The decision to increase the manufacturing capacity depends on the
estimated product reliability.
However, to have the capacity available to meet the expected demand,
the decision has be made and the $5 million committed prior to the
start of production.
Reliability Goal and ALT Discussion
The current product achieves 90% reliability over two years. The best
competitive product is estimated to achieve 98% reliability over the
same period.
The goal for the new design is 99% reliability over two years or better.
This is a major goal and simply conducting ALT is not going to achieve
the result. Yet, a key element lies in understanding whether or not the
goal has been achieved.
The $5 million investment in manufacturing depends on knowing
whether or not the design will meet the goal.
23
finding value – how to determine the value of reliability engineering activities
In this case ALT can answer the question, as it’s focused on the
expected dominant failure mechanism.2
The failure mechanism and the stresses are all known. The new design
using a novel material does leave some uncertainty around how the
design will actually perform. A well-designed ALT program has the
capability to ascertain the expected reliability performance.
ALT Cost
ALT is often expensive to conduct.
The test design, samples, product operation jigs (robots, actuators,
software, etc.), monitoring equipment, and failure analysis all add to
the cost.
Let’s assume the total test planning and setup cost is $50,000.
Demonstrating high reliability will require a significant number of
samples.
The following formula3 provides a rough estimate of the number of
samples needed for a test to demonstrate 99% reliability with 90%
24
Methods to Estimate Accelerated Life Testing Value
confidence under the assumption that no tested samples fail:
ln ^1 - C h ln ^ 1 - 0.9h
ln ^Rh ln ^ 0.99 h
n= = , 230
where
n is the sample size,
C is the statistical confidence, and
R is the reliability.
The 230 sample number is based on a success testing approach in
which the failure mechanism and associated stress are assumed to be
well understood.
Reducing the sample size through the use of degradation testing (or
some other method) may increase the testing complexity but result in
lower overall costs.
The cost of the subsystem that holds the seal is $200 per unit.
25
finding value – how to determine the value of reliability engineering activities
The total cost for samples is then estimated as 230 × $200 = $46,000.
The total cost of the ALT is therefore approximately $96,000.
ALT Value
In this situation the test results provide a binary result.
The population either does or does not achieve at least 99% reliability.
Keeping in mind the ALT is conducted with a sample to represent the
population, there is some uncertainty about the results. Statistical
error may lead to four outcomes, as shown in Table 1.4
Unknown actual
reliability is <99%
Test result Is True Is False
R ≥ 99% Type I error Correct
R < 99% Correct Type II error
Table 1 Statistical errors
26
Methods to Estimate Accelerated Life Testing Value
Assuming the test design used 90% confidence and has a 90% power,
we have a 10% chance of thinking the reliability is less than it actually
is and to opt not to invest in added manufacturing capacity (lost
opportunity for increased sales due false expectation of high field
failure rate then actually occurs).
Further, there is a 10% chance of thinking the reliability is better
than 99% when it is not (a false conclusion), thus investing in added
capacity when demand will not materialize.
Before the ALT we have a 50/50 chance that the new material and
design will meet the 99% reliability goal.
Combining that with the uncertainty of statistical error and a $5
million decision, we can calculate the value of the test.
ALT ROI
The return on investment (ROI) is the ratio of the expected return over
the cost. So, $4 million/$96,000 results in ROI > 41.
Of course, the $5 million decision isn’t the only factor in the value of the
ALT.
27
finding value – how to determine the value of reliability engineering activities
It also provides a baseline for further testing (test cost savings); it may
provide information on the amount of margin the design has over the
goal and permit further design enhancements.
It also confirms the change in reliability, allowing proactive changes in
warranty accruals and service and repair operations.
28
An Estimate of Highly Accelerated Life Testing Value
An Estimate of Highly Accelerated Life
Testing Value
Estimating the value of specific reliability activities is always required.
This is needed to justify the investment required to accomplish the
task. Prototypes, diagnostic equipment, and environmental chambers
are expensive.
One major difficulty is our inability to know what will be found, prior to
conducting the experiment.
Not doing the test means the certainty of not finding anything. Yet, this
lack of knowledge is often not enough motivation to invest in a project
to learn something about the reliability performance.
Highly accelerated life testing (HALT) is a stress testing methodology
for accelerating the discovery of likely failure mechanisms during the
engineering development process.
The following scenario is just one situation. By studying it and
incorporating the few ideas presented will help you estimate the value
of investments in reliability work.
29
finding value – how to determine the value of reliability engineering activities
HALT and Time to Market
Consider the development of a new game controller. The product is a
high-volume seller, with the majority of sales expected immediately
after product launch, during the holiday sales period.
It’s a new design. There’s an emphasis on time to market, with the
majority of the product being manufactured prior to the start of sales.
There are no repairs needed and the controller is an enabling part of a
larger system.
The controller’s reliability goal is 98% reliable over the first year of
ownership when used as part of the game system.
HALT vs ALT
One of the basic questions facing the team is, ‘Will the product meet
the 98% reliability goal?’
ALT may help answer this question, if we know which failure
mechanism(s) will lead to failure during the first year.5
30
An Estimate of Highly Accelerated Life Testing Value
This is a new product without any field history. Other controllers
designed for this environment have experienced a range of failure
causes, but these are often dominated by shock and vibration damage
from dropping.
From the risk analysis done by the design team drop damage is
suspected to be the most significant contributor to product failures.
The new controller is different enough that existing field data are not
likely to be applicable.
Also, it is unknown which specific element of the design would
experience failure first or at all over one year of use. Therefore,
discovering the most likely failure mechanisms that are to occur is
important.
The initial project plan did not include HALT on the first set of
prototypes.
Rather, in the initial plan samples were taken from the second set
of prototypes, 8 weeks later, just before the transfer of the design to
manufacturing to conduct design verification testing (DVT), including
life testing.
31
finding value – how to determine the value of reliability engineering activities
The drop testing portion of the DVT is expected to take a week to
accomplish.
The reliability engineer on this program recommends performing
HALT on the first available prototypes.
The suggestion is to use high loads of random vibration and high shock
loads in the HALT plan to quickly assess the design weakness related to
product drop damage.
The project manager requests more information on timing, cost, and
benefits (value).
HALT Cost
There isn’t time to procure a HALT chamber within the development
schedule; therefore let’s collect quotes from HALT labs to conduct the
testing.
Let’s assume a quote of $10,000 for one round of testing.6
Of course, if there were HALT facilities internally available this cost
would be less.
32
An Estimate of Highly Accelerated Life Testing Value
Also consider that the cost of the prototypes is about five times more
expensive than second-round prototype units.
The first round of prototypes entails a small run, specialized tooling,
and quick turnaround production, ultimately costing approximately
$1,000 for each unit.
Let’s request five units, at an increased cost of five times over later
prototypes: $4,000.
Rounding out the expected costs of engineering support, testing
equipment support, and failure analysis support, we can estimate an
additional cost of approximately $10,000.
Therefore, the total cost to the program to add HALT is approximately
$24,000.
HALT Value
One of the primary benefits of HALT is its potential to uncover new
failure mechanisms in the design.7
33
finding value – how to determine the value of reliability engineering activities
By conducting HALT on the first available prototypes, the design team
increases the time available to resolve design errors or make design
improvements.
Designers tend to design away from failures (design to minimize
field failures); HALT is a tool to discover previously unknown (or
unsuspected) failure mechanisms.
Let’s assume (for the purpose of this example) that the design prior to
any testing has a 25% chance of a failure mechanism that will lead to
an unacceptably high first-year failure rate.
In discussions with the program manager, the team learns that they
would delay the start of production if there were a 10% or higher
expected field failure rate.
Moreover, the cost of the delay was estimated at $500,000 per day in
lost sales.
With an assumed 30 days to design and implement an improvement to
resolve a major reliability issue, the losses would amount to $500,000/
day for 30 days, or $15 million.
34
An Estimate of Highly Accelerated Life Testing Value
There is a good chance that the design is fine and will meet the
reliability objectives.
Let’s assume that 75% of the time the design has an overall failure rate
of less than 10% during the first year.
Also assume that 25% of the time the underlying design has at least
one major failure mechanism that may be detected and resolved prior
to the start of sales.
Consider that no testing program can uncover all faults—yet let’s
assume that only 10% of the time HALT and DVT will not reveal a major
(>10% failure rate) issue. These tests are pretty good at finding major
issues.
Also, let’s say that, in cases where HALT may not reveal the issue,
DVT does identify the fault 50% of the time. And, let’s assume that
HALT identifies the fault only 40% of the time. HALT takes less time to
accomplish and has a higher risk of not finding issues.
Note that this low rate is pessimistic for an estimate of the ability of a
well-executed HALT; in actual practice, HALT is much more effective.
35
finding value – how to determine the value of reliability engineering activities
The value calculation proceeds as follows:
We have a 25% chance of an unacceptable failure rate existing in the
design and a 40% chance of HALT revealing the issue.
We multiply these two chances by the cost avoided by having time to
solve the issue without a 30-day program delay.
This results in an expected savings of 0.25 × 0.40 × $15 million = $1.5
million.
HALT ROI
The ROI is the ratio of the expected return over the cost.
$1.5 million divided by $24,000 gives ROI > 60.
This is only part of the value as we only considered the detection of
major issues, thus avoiding a schedule slip.
HALT will also reveal less significant issues that wouldn’t have resulted
in a schedule slip, yet the earlier detection would reduce the cost of
implementing design changes.
36
An Estimate of Highly Accelerated Life Testing Value
Plus, HALT may have revealed unique failure mechanisms beyond what
the DVT would identify, leading to an incremental reduction in achieved
field failure rate.
37
finding value – how to determine the value of reliability engineering activities
38
Estimating the Value of Derating
Estimating the Value of Derating
This example is based on a real situation.
After a class on design for reliability, a senior manager declared that
every component would be fully derated in every product (electronic
test and measurement devices).
Within a year the design team redesigned all new and existing
products, with strict adherence to the derating guidelines provided in
the class.
A year after the class the product line experienced a 50% reduction in
warranty claims.
They learned about derating and a manager saw the potential value.
We often do not have a manager with such foresight, so we need to
provide justification for the investment.
Here is a case that provides a way to view reliability investments and
determine the return.
39
finding value – how to determine the value of reliability engineering activities
Derating and Field Failure Rate
The specialized test and measurement industry creates very complex
electronic equipment. These are expensive tools with total production
of perhaps 50 per year over.
Like other high cost, low-volume products the cost of failure is very
high.
Because the unit costs are very high, the ability to test sufficient
numbers of units to failure is severely limited.
It is not uncommon to have only one or two units available for all
qualification testing.
Furthermore, the complexity of the units provides multiple possible
failure mechanisms and only rarely does the design provide a clearly
dominant failure mechanism on which to focus reliability evaluations.
Given the barriers to conducting physical testing, the reliability team
recommends implementing detailed derating analysis for the selection
of every electronic component.
40
Estimating the Value of Derating
The design team does use some derating concepts, yet these are only
based on a 50% guideline and no detailed analysis is done.
Therefore, the project manager has requested more information about
the process, costs, and value.
Derating and Field Failures Discussion
Derating is the selection of components that have ratings (power,
voltage, etc.) above the expected stress.8
For example, selecting a capacitor that bridges a 5-volt potential that
has a voltage rating of 10 volts would be considered a 50% derating.
Selecting components that only match the expected stress and rating
generally lead to premature failure of the components, because
the ratings provided by vendors only imply that the component can
experience the stress at the rated value for a very short time.
Derating provides a margin for minimizing the accumulation of
damage or the chance exposure of high enough stress to cause a
failure. You can apply the same concept to mechanical designs, using
safety margins.
41
finding value – how to determine the value of reliability engineering activities
At Hewlett-Packard, a study of the effects of various design for
reliability tools uncovered a very high correlation between well-
executed derating programs and low field failure rates.
This contributed to the 50% fewer field failures experienced.9
In one particular division where the design team embarked on a full
implementation of derating on all products, a 50% reduction in field
failures was achieved in the very first year.
They continued to reduce failure rates over subsequent years as more
fully derated product designs shipped.
Derating Cost
Components that are rated higher cost more and are generally larger
in size.
If the current material cost is $100,000, then implementation of
detailed and thorough derating would likely raise this cost to $200,000.
For a production run of 50 units, the total cost increases by $5 million.
42
Estimating the Value of Derating
The additional engineering time for training, circuit analysis, and
procurement may add an additional $1 million to the project cost.
The total additional cost to the program is thus about $6 million.
Derating Value
The primary value of component derating is the increase in circuit
robustness of the product, which leads to fewer field failures.10
The cost of a field failure is expensive, owing to the replacement cost,
failure analysis, and possible redesign and qualification costs.
Let’s assume that each field failure has an average cost of $2 million
(four times the sales price).
Reducing a 10% annual failure rate (a low estimate for such complex
products) to 5% would result in 2.5 fewer $2 million failures per year
for an annual savings of $5 million.
43
finding value – how to determine the value of reliability engineering activities
Derating ROI
As stated earlier, ROI is the ratio of the expected return over the cost.
With a cost of $6 million and a return of only $5 million, ROI < 0.83.
If the starting failure rate or cost of failure is low, then this ROI may not
exceed the break-even point.
Implementing derating may not make sense in this situation.
Also, consider the market and impact on competition.
If the high failure rate caused a loss of market share, that may further
increase the cost of failure.
Thus consider each situation carefully to ascertain the potential value.
Derating often provides significant value when first implemented,
and if the the team already does a reasonable job of selecting derated
components, the resulting value is diminished.
44
Reliability Prediction Value
Reliability Prediction Value
Reliability prediction entails the forecast or prognostication attempting
to quantify either the time till failure, expected future failure rate or
warranty claims, or required number of spare parts.
One of primary questions confronting reliability engineers is ‘How long
will it last?’
As we make decisions today, we need to know about the reliability of
the design.
We need the information to make decisions so that we can change the
design, stock more spares, or set customer expectations appropriately.
Of course, the best prediction is made after everything has failed.
By shipping all your products and tracking the actual failures over
time, you would know the reliability, but, unfortunately, this only
provides information on what happens after it happens.
45
finding value – how to determine the value of reliability engineering activities
Methods of Prediction and Costs
As Neils Bohr once concluded, “Prediction is very difficult, especially if
it’s about the future.’’
Engineers assemble available information and knowledge and employ
a range of techniques to create reliability predictions.
Methods can range from a simple guess based on engineering
judgment to detailed physics of failure modeling for specific failure
mechanisms and environments.
Engineering Judgment
Early in the concept phase of any design, we may ask, ‘Will this product
last long enough?’
Based on engineering judgment we select basic design elements,
materials, and architecture.
Largely based on engineering judgment, a simple guess helps to
uncover basic information about the use environment and customer
expectations, along with basic technology capabilities.
46
Reliability Prediction Value
Of course, a simple guess has a large degree of uncertainty.
It is however the fastest (i.e., just taking the time for a considered
opinion to be rendered) and least expensive prediction to make.
Parts Count Prediction
In the 1970s and 1980s, many large organizations would diagnose
product failures to the component level. They would keep track
of component-specific failure rates and types of applications or
environments.
These databases of failure information provided a viable means to
estimate future designs.
Large databases such as those of Mil Hdbk 217 or Bellcore (now
Telcordia) collected failure information across broad groups of
products and environments.
Today, most organizations do not conduct detailed failure analysis,
opting to quickly replace or repair the product for the customer, so the
source of information for databases has diminished.
47
finding value – how to determine the value of reliability engineering activities
The basic idea that each component contributes a possible failure and
that adding the individual failure rates gives an estimate of the product
failure rate is full of inaccuracies and faulty assumptions.
Unless one is using an accurate database the results are not much
better than those provided by a simple guess, yet they do assist in
identifying potential reliability issues in a design.
To conduct the study, one needs access to an appropriate failure rate
database, vendor data to supplement for missing values, and a little
time to tally the failure rates.
With a reasonably complete database it may take one or two days to
create a prediction.
Making such a study would involve a small investment but not much
improvement in prediction accuracy would be obtained.
Similar Product History
As organizations stopped conducting detailed failure analysis on every
product failure, they did continue to track product performance.
48
Reliability Prediction Value
Since most products are variations of existing products, the
prediction approach of using similar products as a foundation, then
supplementing with other prediction methods for the new elements, is
a reasonable approach.
For example, if a new design includes the existing power supply of an
existing product, we can use the field failure rate information for that
power supply in the new product.
One must consider the power supply environment and load in the
new design and if any changes in conditions impact its reliability
performance.
This is a good approach to narrow down the areas needing detailed
analysis.
Of course, the approach is only viable if you have the information on
previous products and sufficient failure rate detail for subsystems. For
completely new designs, this is not possible.
The use of similar product history is comparable to parts count
prediction and generally has better accuracy. Like a parts count
approach it may take one or two days to track down and tally the
prediction.
49
finding value – how to determine the value of reliability engineering activities
Weakest Element Estimate
A product will often fail as a result of its weakest element failing, so
estimating this element’s failure rate can be a useful prediction.
Let’s consider an organization’s tape backup system. Like any
electromechanical system it has many possible failure mechanisms.
If it is assembled, transported, and installed correctly the dominant
failure mechanism is the read/write head wear caused by tape
abrasion.
When the head dimension is reduced enough, the ability of the device to
read/write ceases.
Each length of tape dragged across the head creates a predictable
amount of wear; the organization accurately predicted the time to
failure based on amount of use (i.e., feet of tape passing across the
head).
Field data on similar models verified that the principal reason for
product failure was head wear.
50
Reliability Prediction Value
The design team optimized the wear and performance, yet read/right
head wear remained the primary failure mechanism for the product.
In this kind of situation, the team has the benefit on being able to focus
on one mechanism and creating a product prediction.
The other elements of the system only had to last longer than the head.
This approach to prediction simplifies the amount of work and
experimentation and provides an accurate life estimate.
Of course, changes to the materials, tape, speed, tension, and other
variables affecting the head wear will change the relationship.
Moreover, the team must maintain due diligence with the reliability of
other elements to avoid creating a new weakest link.
Reliability Block Diagrams
Block diagrams are an organizational tool for considering
contributions to the system reliability.
51
finding value – how to determine the value of reliability engineering activities
Like an organization chart, each block is a subsystem or element of a
product.
Let’s say we have a desktop computer as the product.
The top block is the system; below it, there are five blocks that
represent the power supply, hard drive, mother board, display, and
keyboard.
Block diagram structures exist for series, parallel, and complex
systems. Each block includes the reliability of that element.
Depending on the structure (reliability-wise) the calculations for
determining the system reliability may differ.
52
Reliability Prediction Value
Although constructing a block is not a prediction method on its own,
once one is established, the ability to compare design options (say, two
hard drives with different expected reliability performance) and the
impact on the system reliability becomes easier.
In one regard it resembles the parts count method with the added
benefit of being able to account for parallel reliability structures.
Experimental Results
Life testing takes many forms and it is beyond the scope of this book to
describe them all. You can find many books on accelerated life testing
to provide detailed guidance.
As a prediction method, conducting an experiment is expensive, but
the results are accurate when the experiment is done well. The basic
requirement is knowing the failure mechanisms of interest and how the
appropriate stress relates from experimental to use conditions.
Errors or poor assumptions can make this method very inaccurate,
so take care when designing, conducting, and analyzing reliability
experiments or tests.
53
finding value – how to determine the value of reliability engineering activities
Physics-of-Failure Modeling
Research and modeling have enabled detailed characterization of
failure mechanisms.
Although not every possible failure mechanism has a solid physics-
based model, many are well known and have a physics-of-failure
model.
Simple models include solder joint fatigue formulas based on the
Coffin-Mason relationship.
Complex models may require finite element tools.
The technical literature provides details on a variety of physics-of-
failure models.
If a model doesn’t exist you may need to conduct experiments to fully
characterize the relationship between stress and life performance.
Once you establish a physics-of-failure model, you have the ability to
consider changes in the environment, stress load, structure, material
set, or dimensions (variables affecting the failure mechanism) to
create reliability predictions.
54
Reliability Prediction Value
Physics-of-failure models take time and are expensive to create but
they provide the most flexibility and offer higher accuracy than other
methods.
Sources of Value
We create reliability predictions to support decisions. Predictions can
help us answer the following questions:
• Is the product reliable enough?
• Is the product going to meet our reliability targets?
• How many spares will we need over the next year?
Many other specific questions can be addressed when we’re not able to
wait for actual results to materialize.
Making Informed Decisions
Let us consider the simple example of comparing vendors of hard
drives.
Understanding an accurate reliability prediction for our application
55
finding value – how to determine the value of reliability engineering activities
enables selecting the most cost effective and reliable product to meet
our business goals.
Without a prediction, we may select a hard drive that fails too early and
limits the product’s reliability—or we may select one that is too reliable
and expensive for our application.
Identifying and Improving Design Weaknesses
During product or system design, the ability to identify the elements
that will lead to failures (weakest links) enables us to focus resources
on those areas for improvement.
Without knowing where to focus we may have more failures than
anticipated along with the remorse of ‘if we only knew.’
Allocating Resources Appropriately
Beyond focusing efforts for reliability improvements, predictions also
divert resources away from elements that are very reliable already.
56
Reliability Prediction Value
Prediction also focuses product testing on the failure mechanisms
most likely to limit product reliability. This may help reduce the costs of
prototypes and testing facilities.
Using reliability block diagrams and similar models enables the
balancing of reliability with component costs to optimize the reliability
at the minimum cost.
Setting Expectations
An accurate prediction is useful inside the company to forecast
warranty and repair costs.
Outside the company, predictions are useful to customers when:
• making purchasing decisions,
• planning larger systems using the product reliability information for
modeling, and
• estimating total cost of ownership, including spares, downtime, and
maintenance costs.
The value of reliability predictions in in the customer understanding
57
finding value – how to determine the value of reliability engineering activities
of the claims. Product data-sheets, reliability white-papers, warranty
policies all contribute to the reliability contribution to the brand.
Finally, creating a prediction and comparing it to field reliability
performance allows the organization to improve its next-generation
product by using the field data for similar products and refining any
models created for reliability predictions.
58
Increasing Value
Increasing Value
Value for any business activity serves as a guiding principle for staying
in business.
Running a business if not always about money—it is about value.
Businesses invest in product design and distribution and plant design
and operation to realize a benefit or return on the investment that
makes the investment worth the effort and time.
Projects may create value by increasing one or more of the following:11
• revenue,
• profit,
• growth,
• offerings,
• retention,
• return on investment,
• return on assets,
• efficiency,
• visibility,
• equity, and
• net preference.
59
finding value – how to determine the value of reliability engineering activities
Reliability improvements may impact all of these.
Focusing on one aspect to provide the value, and letting others
contribute as they may, provides the means to create the ROI
justification for a proposed reliability improvement project.
The easiest to understand is profit.
If you offer a warranty or incur costs when a product or asset fails,
then you can increase profit by reducing failures, that is, by making the
product or equipment more robust such that it works longer.
Increased visibility can translate into sales, leads, opportunities, or
what other benefits. Of course, here we are talking about the positive
benefits of visibility—not the news headlines of court indictments,
illegal insider trading, or other types of notoriety.
Creating a product that creates word-of-mouth recommendations is
fundamental to success.
For example, consider the last time you rented a car. Was it a model
that you would recommend? .
60
Increasing Value
As we move though life we experience products and let those around us
know how well they work.
As you create products one of the key attributes is the reliability of the
product: If you let down your customers, think of the stories they will
tell.
61
finding value – how to determine the value of reliability engineering activities
62
11 Ways to Find Reliability Value
11 Ways to Find Reliability Value
Most everyone agrees that improving a product or process reliability is
a good thing. It’s good for customers, factories, and our business.
But sometimes it’s difficult to answer the question, ‘What is the value
of that reliability activity?’
What if your boss asks you what value you provide to the organization.
Your answer may be harder to compose than you think.
How would you quantify your skills, experience, and knowledge
and your role within the complex formal and informal working
environments?
Here’s a list of ways to uncover the value in your reliability program.
You can use this as a way to show potential value for future projects or
to capture and record value from past activities.
Either way, this list provides a sound basis for planning and focusing on
adding value to your organization through reliability engineering.
63
finding value – how to determine the value of reliability engineering activities
Cost Reduction
Cost reduction should be rather straightforward to calculate.
Many consider only the cost of components for cost reduction. But this
sort of reduction often increases failure rates and the costs of failures.
So, consider reliability improvements and the change in expected
failure rate for the cost reduction element.
Indeed, you may need to increase component costs to improve
reliability, but doing so is warranted if the improved reliability reduces
the cost of failures.
Another way to achieve cost reduction is to investigate design or
process faults earlier in the life-cycle.
The basic idea is that the cost of resolving a design issue can increase
by an order of magnitude per stage along the life-cycle.12
For example, if it costs the design team $100 to resolve an issue
during the concept phase, it may cost $10,000 to resolve it during the
production phase (two stages later).
64
11 Ways to Find Reliability Value
Therefore, if the reliability work assists in identifying and fixing issues
earlier than later, you are avoiding the costs of fixing the same issue
later.
Warranty Reduction
Another obvious benefit of improving product reliability is that
reducing the number of failures that customers experience leads to
fewer warranty claims.
To find the cost per failure for your product or system, gather the cost
of warranty or the cost of failures and divide by the number of failures
(warranty claims) experienced.
If you avoid 10 future failures then that represents a savings of 10 times
the cost per failure.
Another way to track and claim your warranty impact is to divide the
cost of warranty by the total number of products sold. This provides the
cost of warranty per unit sold. This is the same cost basis as used for
individual components.
65
finding value – how to determine the value of reliability engineering activities
For production equipment, for which cost per day to operate the plant
is the vital metric, divide by the number of days to get cost of repair
(downtime) per day. If the vital metric is cost per unit produced, then
divide by number of units.
Risk Reduction
Risk is uncertainty.
• Will the product perform as expected?
• Will the equipment have the rated capacity?
• Have any major faults been overlooked?
• How will the customers actually use this device?
For any product in design there are many other unknowns that provide
risk.
We can quantify risk by asking a couple of questions:
1. Does the reliability work mitigate or reduce field-related problems?
If so, then we can estimate the probable cost of the field problem in
dollars (i.e., units affected times the repair cost).
66
11 Ways to Find Reliability Value
2. Has the probability of field-related problems been reduced?
If so, then we can give a estimate by how much (e.g.,, an estimated
1,000 units per month with a $50 cost per failure with a reduced risk of
5% leads to a value of $2,500 per month).
Time to Market Impact
If an organizational objective is related to time to market, then there
may be a substantial value in minimizing risks to the development time
line.
The discovery of major reliability issues late in the process can delay
a program launch. Identifying issues earlier provides less-expensive
means to resolve them, and reduces the risk of delaying the program.
To find the value consider the following thought process:
• Did the work identify any problems with potential impact to time to
market (TTM)?
• Has the use of tools or techniques identified issues that may impact
TTM?
67
finding value – how to determine the value of reliability engineering activities
If the above apply, then
• Identify the types of problems.
• Estimate the cost of delay in TTM.
• Determine the opportunity in dollars of additional income from an
early TTM.
Another factor to consider is the additional cost of engineering or
development teams during the extended or reduced TTM.
Time to Volume Impact
Related to time to market is time to volume. For high-volume products
the ramp-up to production levels may be an important element of the
overall business plan.
When there are unknowns or major risks the ramp-up of production
may slow to minimize the risk.
By minimizing the risk to production (thereby increasing confidence
that production is making good products) the ability of the team to ship
adequate numbers of units for additional markets increases the impact
of initial product offering marketing.
68
11 Ways to Find Reliability Value
Other relevant questions to ask are these:
• Did the work help the team accelerate or meet your time to volume
(TTV) goals?
• If applicable, what is the estimated dollar impact of avoiding the TTV
issues that were resolved?
Material Cost Reduction
Material cost reduction involves the cost of yield loss during
production, the cost of scrapping bad batches of incoming or outgoing
material, and the cost of recalls.
The cost of prototypes or testing materials must also be considered.
The key question is whether any direct product material or test
equipment costs can be avoided or reduced.
For example, did an FMEA study identify and resolve a failure
mechanism that would otherwise require ALT to estimate field failure
rates? If so, the FMEA created the value of avoiding the cost of the ALT
equipment and samples.
69
finding value – how to determine the value of reliability engineering activities
Customer Satisfaction Improvement
For many products there is a direct relationship between customer
satisfaction and product reliability. Satisfied customers buy more
products and encourage others to do so. Dissatisfied customers do not
buy products and may discourage others from buying.
Another source of value has to do with support: If the product reliability
is improved, customer call centers and repair centers do not need
as large a staff or facilities. Some organizations identify the cost per
customer call, which allows an estimation of value of a change in call
rates.
Consider whether the reliability work impacts customer satisfaction
and, if so, how and to what extent. For example, determine how many
customer calls would be avoided.
If you have a model of how customer satisfaction affects sales, you can
estimate the impact to sales volume.
Often the revenue value is significant. Although gathering good
numbers and developing models for these estimates can be difficult to
create, they are well worth the effort.
70
11 Ways to Find Reliability Value
Opportunity Costs Reduction
Every engineer and manager on a development team has multiple
priorities and tasks to accomplish. If members of the team are diverted
to accomplish a reliability task, they are not performing their primary
role.
If the means by which a task is accomplished can be automated,
streamlined, or out-sourced effectively, this reduces the lost
opportunity for engineering tasks to be accomplished.
A clear example is to consider the cost of extending the program
development time and the cost per day of the development team.
If a reliability activity can reduce the time of the extension, this saves
the opportunity cost and those engineers can then work on the next
project as planned.
Indirect Impact
The indirect impact is more difficult to quantify.
Consider the following:
71
finding value – how to determine the value of reliability engineering activities
• Did the reliability activity increase the efficiency of the team?
• Did the activities bring previously unknown knowledge to the
program?
• Did the work improve the team’s ability to make decisions with fewer
errors?
Engineering Effort Saved
Consider as an example a problem with cracks developing in a line of
capacitors. The reliability engineer is asked to research and resolve
this cracked-capacitor issue. It may take about two weeks to conduct
the research and resolve the issue.
The value was in part just in the engineering work, but the additional
and significant value for the organization was in the reuse of that two
weeks of work.
Over the next several months other similar situations might more
easily be resolved, saving about two weeks each time.
Let’s assume that over the next 3 months, another 12 similar situations
arose.
72
11 Ways to Find Reliability Value
We can estimate the total engineering effort saved by
(cost of an engineer per day) × (2 weeks) × 12
representing a savings of half a man year of engineering time.
This savings was in addition to avoided material costs, field failure
costs, and impact to TTM or impact to customer satisfaction.
Final Thoughts
In any organization or market you can identify meaningful sources of
value.
It takes practice, but over time you will naturally look for and quantify
value for each of your reliability-related activities. As you do so, you
will be able to clearly estimate return on investment for proposed tasks
and track and qualify value created as a result of specific tasks.
In a large part, it’s about learning to talk as your management team
does about investments, opportunities, and profits. Doing so increases
your influence within an organization and improves the product
reliability for your customers.
73
finding value – how to determine the value of reliability engineering activities
74
Notes
Notes
1. Patrick D. T. O’Connor and Andre Kleyner, Practical Reliability
Engineering, 5th ed. (Chichester, England: John Wiley & Sons, 2012) 1.
2. Wayne Nelson, Accelerated Testing: Statistical Models, Test Plans,
and Data Analysis (Chichester, England: John Wiley & Sons, 1990) 3.
3. Gary S. Wasserman, Reliability Verification, Testing and Analysis
in Engineering Design (New York: Marcel Dekker, 2003) 209.
4. R. Lyman Ott, An Introduction to Statistical Methods and Data
Analysis (Belmont, CA: Duxbury, 1993) 216.
5. Mike Silverman, How Reliable Is Your Product? (Cupertino, CA:
Super Star, 2010) 193.
6. Mike Silverman, personal communication, discussion about
average cost of HALT not including the prototype costs, June 18, 2011.
7. Gregg K. Hobbs, Accelerated Reliability Engineering: HALT and
HASS (Chichester, England: John Wiley & Sons, 2000) 43.
8. W. Grant Ireson, Clyde F. Coombs, and Richard Y. Moss,
75
finding value – how to determine the value of reliability engineering activities
Handbook of Reliability Engineering and Management (New York:
McGraw Hill, 1995) 16.9.
9. Richard Y. Moss, personal communication discussion his
experience teaching and tracking results concerning derating, June 12,
2002.
10. W. Grant Ireson, Clyde F. Coombs, and Richard Y. Moss,
Handbook of Reliability Engineering and Management (New York:
McGraw Hill, 1995) p. 5.4.
11. Fields, D. The Executive’s Guide to Consultants: How to Find, Hire
and Get Great Results from Outside Experts, Kindle ed. (New York:
McGraw-Hill, 2012) Table 1-5.
12. Jonette M. Stecklein, et al., 2004, “Error Cost Escalation
Through the Project Life Cycle”, Paper presented at the 14th Annual
International Symposium of the International Council on Systems
Engineering (INCOSE) Foundation (Toulouse, France, June 19, 2004).
76
Are you Ready to Accelerate
your Reliability Program and Career?
We’ve put together a comprehensive remote support and
mentoring program, which we call Reliability Coaching.
The book you’ve just read covers one element of creating an
effective reliability program or career … and that’s only the
beginning.
We’ve been working on hundreds of projects developing products,
streamlining maintenance, and improving reliability programs
for over 20 years. We’ve been fortunate to enjoy a lot of success in
that time, and it took a lot of work … and we’ve made our share of
mistakes along the way.
What if you could directly benefit from those years of experience—
and avoid those mistakes?
What if you could easily learn and apply reliability engineering best
practices, tools, and resources?
What if you could create a culture of reliability in your organization
with everyone working toward the same goals?
We’ve got something to show you. We call it Reliability Coaching,
and it’s the best way to enhance your reliability program & career.
[Link]/reliability-coaching/
Finding Value
How to Determine the Value of Reliability Engineering Activities
Fred Schenkelberg
Fred Schenkelberg is an international authority on reliability
engineering. He is the reliability expert at FMS Reliability, a
reliability engineering and management consulting firm he founded
in 2004. Fred left Hewlett Packard (HP)’s Reliability Team where he
helped create a culture of reliability across the corporation to assist
other organizations. His passion is working with teams to improve
product reliability, customer satisfaction, and efficiencies in product
development; and to reduce product risk and warranty costs. Fred’s areas of expertise are:
reliability program development, accelerated life test design and analysis, reliability statistics,
risk assessment, test planning, and training. He has a Bachelor of Science in Physics from
the United States Military Academy and a Master of Science in Statistics from Stanford
University.
About this book
• Estimating or calculating value of any activity enables you to:
• Make convincing proposals
• Select the best approach
• Demonstrate your worth
• Improve you influence
This short book details 11 ways to uncover and determine the value of a wide
range of reliability activities. You will understand the importance speaking in
terms of value. There are detailed examples to show you the thought process of
a practical approach to finding value.
Design : Product : Management
& Leadership : Quality Control
ebook ISBN: 978-1-938122-02-6
paperback ISBN: 978-1-938122-03-3