0% found this document useful (0 votes)
60 views17 pages

Availability Calculation in Reliability

Reliability, availability, and maintainability are key metrics used to define and quantify dependability. Reliability is measured by mean time to failure (MTTF) and failures in time (FIT), while availability considers up and down time and is calculated as MTTF/(MTTF + mean time to repair (MTTR)). System reliability models like triple modular redundancy (TMR) can improve reliability over simple redundancy and is calculated based on the individual component reliabilities. The "bathtub curve" illustrates how failure rates change over the lifetime of a component.

Uploaded by

anon_240227768
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
60 views17 pages

Availability Calculation in Reliability

Reliability, availability, and maintainability are key metrics used to define and quantify dependability. Reliability is measured by mean time to failure (MTTF) and failures in time (FIT), while availability considers up and down time and is calculated as MTTF/(MTTF + mean time to repair (MTTR)). System reliability models like triple modular redundancy (TMR) can improve reliability over simple redundancy and is calculated based on the individual component reliabilities. The "bathtub curve" illustrates how failure rates change over the lifetime of a component.

Uploaded by

anon_240227768
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

CS203 – Advanced

Computer Architecture
Dependability & Reliability
Failures in Chips
Transient failures (or soft errors)
Charge q = c*v if c and v decrease then it is easier to flip a bit
Sources are cosmic rays and alpha particles and electrical noise
Device is still operational but value has been corrupted
Intermittent/temporary failures
Last longer
Due to
Temporary: environmental variations (eg, temperature)
Intermittent: aging
Permanent failures
Means that the device will never function again
Must be isolated and replaced by spare

Process variations increase the probability of failures

2
Define and quantify dependability
Reliability =
measure of continuous service accomplishment (or time to failure).
Metrics
Mean Time To Failure (MTTF) measures reliability
Failures In Time (FIT) = 1/MTTF, the rate of failures
Traditionally reported as failures per 109 hours of operation
Ex. MTTF = 1,000,000 FIT = 109/106 = 1000
Mean Time To Repair (MTTR) measures Service Interruption
Mean Time Between Failures (MTBF) = MTTF+MTTR

3
Define and quantify dependability
Availability =
measures service as alternate between the 2 states of
accomplishment and interruption (number between 0 and 1, e.g.
0.9)
Module availability = MTTF / ( MTTF + MTTR)

4
Fault-Tolerance
How to measure a system’s ability to tolerate
faults?
Reliability = Probability[no failure @ time t] = R(t)
Availability = Probability[system operational]
E.g. AT&T ESS-1, one of the first computer-controlled
telephone exchange (deployed in 1960s) was designed for
less than two hours of downtime over its lifetime: 40 years.
Availability = 99.9994%
Failure rate
Fraction of samples that fail per unit time
Is NOT constant, changes over time
R(t) = N(t)/N(0), where N(t) is the number of
operational units at time t.

5
Example calculating reliability
If modules have exponentially distributed lifetimes (age
of module does not affect probability of failure),
Overall failure rate is the sum of failure rates of all the modules
Calculate FIT and MTTF for 10 disks (1M hour MTTF per disk), 1
disk controller (0.5M hour MTTF), and 1 power supply (0.2M hour
MTTF):

FailureRate = 10 ´ (1/1,000,000) +1/500,000 +1/200,000


= (10 + 2 + 5) /1,000,000
= 17 /1,000,000
= 17,000FIT
MTTF= 1,000,000,000 /17,000
» 59,000hours
6
The “Bathtub” Curve
Failure Rate

1 3

Early Life Wear-Out


Region Constant Failure Rate Region
Region

0 Time t
7
The “Bathtub” Curve

Burn-in is a test performed to screen or


eliminate marginal components with inherent
Failure Rate

1
defects or defects resulting from manufacturing
process.

Early Life
Region

0 Time t
8
The “Bathtub” Curve
An important assumption for effective maintenance is that
components will eventually have an Increasing Failure Rate.
Maintenance can return the component to the Constant Failure
Region.
Failure Rate

Constant Failure Rate


Region

0 Time t
9
The “Bathtub” Curve
Components will eventually enter the Wear-
Out Region where the Failure Rate
increases, even with an effective
Failure Rate

Maintenance Program. You need to be able 3


to detect the onset of Terminal Mortality

Wear-Out
Region

0 Time t
10
Derivation of R(t)
Probability[no failure @ time t] = R(t)

Assuming a constant failure rate λ, N is the number of units

dN = -l N (t)dt
dN(t)
dR(t) =
dN(0) Integrating with R(0) = 1 boundary:
R(t) = e-λt
dR(t)
= -l R(t)
dt

11
System Reliability
Series system Parallel system

R1

R1 R2 Rn R2

n
RS = Õ Ri Rn

i=1 n
RP =1- Õ (1- Ri )
i=1

12
Triple Modular Redundancy
TMR: Triple Modular Redundancy
three concurrent devices plus a voter (assume no voter failure)
RTMR(t) = R3(t) + 3R2(t)(1 – R(t)) = 3R2(t) – 2R3(t)
Let R(t) = e-λt, then RTMR = 3e-2λt – 2e-3λt

Voter Result

13
Simplex v/s TMR Reliability
1.0

0.9 TMR has higher reliability Simplex


TMR
0.8
for short mission times

0.7
Reliability

0.6
After 1st failure,
TMR equivalent to
Reliability

0.5 2 component in series


0.4 Simplex

-lt
0.3 RSimplex(t) = e

-2lt -3lt
0.2 RTMR(t) = 3e - 2e

0.1
TMR
0.0

0 1 2 λt 3 4 5

lt 14
MTTF - Mean-Time To Failure
Let F(t) = 1 – R(t), the failure probability (cdf)
and f(t) = dF(t)/dt, the failure probability density
¥
MTTF = ò tf (t)dt
0
¥

ò tle - lt 1
MTTF = dt =
0 l
Expected working life of a unit with
an exponentially distributed
reliability is the inverse of its failure
rate
15
MTBF
The MTBF is widely used as the
measurement of equipment's reliability and
performance.
This value is often calculated by dividing the
total operating time of the units by the total
number of failures encountered.
This metric is valid only when the data is
exponentially distributed.
This is a poor assumption which implies that
the failure rate is constant if it is used as the
sole measure of equipment's reliability.
16
Summary
How to define dependability
How to quantify dependability
How to measure Reliability of a system

17

You might also like