Reliability
Glen Dobson
[Link]@[Link]
[Link]
Recapping
Overview of dependability
the property of a system such that we can justifiably place our reliance on the
service it delivers
Key dependability attributes
Reliability, Availability, Safety, Security
Relationship between attributes
Effect of primary attributes on each other as well as the effect of auxiliary
attributes
Criticality and conflict
Different attributes are critical to different systems
Improving one attribute may be detrimental to another
Dependability requirements
Use measurable criteria
High dependability => Hard (& expensive) to test
Availability
Readiness for correct service
Problems inherent vs. operational availability, what availability does not tell you
Overview
Definition of reliability
Reliability metrics
System failure
Preventing failure
Testing for failures
Group Discussion
Definition
Laprie
Reliability is the continuity of correct service
(a service is correct when it implements the
system function)
More pragmatically
In a given time period, for a given usage
pattern, how likely is a system to fail?
Failure = deviating from the system specification
For some systems Failure may mean deviating from
expectations
Assessment qualifiers
Assessment of reliability depends on:
Intended system usage
Intended operational profile
Context and environment of use
Time and period of use
Load and intensity of use
Reliability is a function of these factors
If any change, we must reassess
Reliability measures
POFOD - Prob. of failure on demand
ROCOF - Rate of occurrence of failure
MTTF - Mean time to failure
Each is suitable for different systems
Suitable time units should be chosen
Physical or logical
Suitable time units???
Determine suitable time units for:
ATM cash withdrawal
Editing with a word processor
Web server providing pages
General practitioner diagnosis
Nuclear reactor core shutdown Alert
Patient illness
Request
Minute
Transaction
Bath tub curve
Time
Failures
Effects:
Hardware (degradation)
Software (evolution)
People (mental faculties)
Software & the bath tub
The bath tub is a sketch of hardware reliability
Tends not to apply so well to software
Burn in period is similar
Upgrades cause a sudden decrease in reliability
Usually followed by another burn in period
Ideally the upgrade/burn in effect decreases over time
Once the software is no longer upgraded then the
reliability becomes constant
Our bath tub will certainly have a bumpy bottom
Software Reliability Curve
Time
Failures
v1.0 v2.0 v3.0
Initial
Development
Software no longer
actively maintained
But these are only sketches
Systems are made up of software, hardware
and what about people?
What does their failure curve look like?
Burn in period = training/familiarisation
May forget/pick up bad habits over time
Organisational changes will have an effect
e.g. high workload/stress will affect mental capacity
Personnel changes will have an effect
So the bath tub is likely to have bumps all over
the place.
Failure manifestation
Fault Failure Error
Fault
The adjudged or hypothesised cause of an error. Typically a
mistake or lack in the preparation of a component.
Error
The initial deviation in system state which eventually leads to
failure. This is usually unintended/unexpected behaviour.
Failure
A deviation from correct service (i.e. from the specification or system
function)
Metaphor
Fault Error Failure
Examples
Fault
Programming mistake
Poor training
Joined pins on chip
Error
Incorrect floating point calculation
Misfiling of documents
No actuator signal
Failure
Mis-navigation
Treatment not given to patient
Burglar alarm not sounding
Fault classification
Can classify along many axes:
Phase of creation/occurrence
Development Faults/Operation Faults
System Boundaries
Internal Faults/External Faults
Phenomonological Causes
Natural Faults/Human-made Faults
Dimension
Hardware Faults/Software Faults
Objective
Malicious Faults/Non-Malicious Faults
Fault classification (2)
Can classify along many axes:
Intent
Deliberate Faults/Non-deliberate Faults
Capability
Accidental Fault/Incompetence Faults
Persistence
Permanent Faults/Transient Faults
Failure classification
Again several axes of classification:
Failure Domain
Content failures/Timing Failures
Or if both halt failure/erratic failure
Detectability
Signalled Failure/Unsignalled Failure
Consistency
Consistent Failures/Inconsistent (Byzantine)
Failures
Consequences
Minor Failures/Catastrophic Failures
Fault-Error-Failure
A Common source of confusion results
from perspective on system, eg..
mental human fault
programming error
software fault
software error
system failure
The general case
Failure
Fault
Error
Failure
Fault Error Failure
Fault
Latencies
Fault Failure Error
Fault latency Failure latency
Faults may go undetected for a long time
Faults may be dormant (i.e. never lead to an error) or
active
Internal errors may never reach the systems external
state
Statistical reliability testing
Identify
Op. Profiles
Test and
Log System
Create
Test Data
Compute
Reliability
Hard to identify operational profiles:
No such thing as standard usage
Exceptional system usage and Maverick users
Patterns of usage change over time
Hard to perform significant number of tests for
very high reliability systems
Automated statistical testing
Capture of operational profiles
Auto generation of minimal test sets
Used to assess the reliability of system
Still unrealistic for VHR systems
Failure prevention
Fault Failure Error
Fault
avoidance
Fault
tolerance
It is better to avoid faults than tolerate them
(prevention is better than the cure)
We might not regain our balance !
Fault
removal
Fault avoidance - Use tarmac instead
Fault removal - Council fix slab, walk around slab
Fault tolerance - Trip, but regain balance
Fault avoidance
Prevent inclusion of new faults
Managed development
Formal methods
Quality culture
Managed development
Mature lifecycle (Certified Process?)
Management and control of:
Requirements and designs
Evolution
Testing
Configuration
Documentation
Traceability
Accountability
Audit and review
Formal methods
Specify system using formal language
Precise vocabulary, syntax and semantics
Based on maths, set theory, logic etc.
Spec. can then be processed formally
Benefits of formal methods
Reduce ambiguity & misunderstanding
Automatically analyse for:
Consistency
Correctness
Completeness
Specs can be emulated or simulated
Verify produced system using proofs
Prove various properties of system
Transformation to construct system
Propositional calculus
Based on sentences and propositions
Allows implies, not, and, or operators
Basic expression and proofs
Predicate calculus
Extends propositional calculus
Much more powerful
Allows the use of variables
for all (universal) quantifier
there exists (existential) quantifier
Popular methods
OBJ - Object oriented, executable lang
VDM - Based on state and operations
Z - Set theory and graphical schemas
Lotos - Parallel & concurrent systems
CCS - Parallel & concurrent systems
Formal method problems
Time consuming and expensive
Hard to understand (not fun)
Domain experts stand little chance
Problems concealed by formality
Transformation slow and difficult
Tool support is patchy
How do we know specification is right?
No single language suitable for everything
Applying formal methods
Useful for specific sub-systems
Useful for specific sub-problems (e.g.
safety)
Cost effective if used appropriately
Limited use in industry
Has yet to deliver in large scale
Fault removal
Fault Failure Error
Fault
avoidance
Fault
removal
Fault removal
Detect and remove existing faults
Testing and debugging
Reviews/inspections
Static analysis
Testing
Alpha/Beta - Acceptance/real operational use
Black/White box - opaque/transparent components
Functional/Structural - (as above)
Defect/Statistical - explicit search/normal usage
Unit/integration - component/whole system
Regression - repeat test set after each repair
Stress - push upper bounds, try to break system
Reviews and inspections
Focus on artifacts produced
No operational system required
Expert judgement - cross discipline
Examine & critique produced artifacts
Requires knowledge of artifacts and domain
Often cheaper than testing
Not all problems are identified
Used to assess non testable attributes
Fault tolerance
Fault Failure Error
Fault
avoidance
Fault
tolerance
Fault
removal
Fault tolerance
Handle faults and resulting errors
Prevent propagation to failure
Fault
Detection
Fault
Recovery
Damage
Assessment
Fault
Repair
Self checking
External monitors
Redundancy
Safe state
Repair
Restore
Checksums
Redundant links
Assertions
Run time checks
Performed periodically
Ensures system in a safe state
Are we safe before we continue?
Manually coded or auto-generated
Modular Redundancy
Triple modular redundancy (TMR)
3 components do same task at same time
Output comparator
Majority voting
Failure likelihood
All modules failing together unlikely
Provided modules are independent !!!
Scaled up to N-version systems
Automatic module repair/replacement
Recovery Blocks
Redundant components used in series
Acceptance test used to assess results
If one component fails, try the next
Roll-back state before retry
Try until success or no more left
Comparing approaches
Modular redundancy less efficient
Since all modules MUST be executed
Recovery blocks good (with no failure)
But how to write the acceptance test ?
Component diversity
Diversity essential for these methods
Both in design and implementation
Each component should use different:
System specifications
Design paradigms
Programming languages
Development environments
Algorithms
Backgrounds and cultures
Problems with redundancy
Duplicate faults can still exist !!!
People still make the same mistakes
Hard to think of different ways to work
Added complexity can hide faults
Cant do acceptance test for everything
What happens if components dont agree?
Big efficiency hit (problem for RT systems)
Can be very expensive (three times the cost)
? Reliability of this course ?
Fault avoidance
(almost formal methods ;o)
Fault avoidance
Redundancy, Diversity
Fault removal
Fault removal
Testing
Redundancy
Diversity
Fault removal - review
Fault tolerance
Trying to avoid the failure of this course:
Used Ians book
Also used alternative sources
Used spell checker to make slides
5th year we have done this course
Both Mark and Glen are lecturing
Checked each others slides
Mark (social sci.) Glen (comp sci.)
Your comments in lectures
Fault injection
Assess system or sub-component
Test harness to assess fault tolerance
Artificially create and introduce faults
Can be performed on:
Simulation of component
Actual component under test load
Actual component in actual use
Each has its own pros and cons
Uses of fault injection
Identify dependability bottle necks
Study behaviour in presence of faults
Assess error detection mechanisms
Assess error recover mechanisms
Assess error repair mechanisms
Fault injection
Fault
Injector
Data
Collector
Workload
Generator
Fault
Library
Workload
Library
Collected
Data
Target System
Controller
Types of injection
Compile time injection
Run time injection
Interactive injection
Addition of new
Alteration of existing
Removal of old
Problems with fault injection
Unrealistic operational profiles
Time consuming for complex systems
Impractical for VHR systems
Instruments interfere with operation
Limited to S/W and H/W components
Or is it ?
Summary
Group discussion
Dave works in an office. He uses a desktop PC with an off-the-shelf
OS. He does most of his work using a standard office suite (word
processor, spreadsheet, etc.). Dave browses the web and e-mails
his friends when he is bored. His dog is called Caruthers.
Dave finds using his computer very unreliable and regularly loses
work. What reasons could there be for this unreliability? What could
be done to improve reliability?
Consider both social and technical perspectives. Do not limit your
thinking to only those topics covered in the lecture material.
Thoughts - problems
Is it unreliable ? From his perspective - yes
Old or faulty hardware (e.g. memory)
Dodgy software
Lack of latest updates
Inadequate training of Dave
Dog hair inside floppy disks
Inappropriate user interface (OS and suite)
Unanticipated use of suite (Excel games)
Virus infection from e-mail
Unreliable auxiliary s/w (e.g. browser)
Thoughts - solutions
Hardware checking and repair
Install latest s/w updates
Proactive behaviour by Dave (regular saving)
Backing up procedures
Avoiding unreliable features (e.g. tables)
Redundancy - Davina replicates work
Better testing, reviews, walkthroughs etc
Formal modelling of Daves office (NOT)
Ethno study of Daves office (insight ?)
Not cost effective to provide high reliability !
Further questions
What effect would the publishing of
portions of the OS on the web have?
What effect would a new competitor for the
office suite have?
What effect would the installation of a new
version of the OS have?