Module 6.
1: Maintenance Task Selection
Learning Objective: To master the technical criteria for selecting condition-
based, age-related, and failure-finding tasks, with a specific focus on the critical
concepts of the P-F Interval, useful life, and MTBF, and how they dictate task
applicability and frequency.
o Condition-based, Age-related, and Failure-finding Tasks
This section refines the task selection hierarchy with precise technical
definitions and applicability criteria.
1. Condition-Based Tasks (CBT) / Predictive Maintenance (PdM)
Definition: A scheduled task to monitor a specific physical parameter
that indicates a potential failure is in progress.
Objective: To detect the Potential Failure (P) point so that corrective
action can be scheduled before the Functional Failure (F) occurs.
Technical Applicability Criteria: A CBT is applicable ONLY if all of the
following are true:
1. It is possible to detect reduced resistance to failure (the
potential failure condition).
2. The P-F interval is reasonably consistent.
3. The P-F interval is long enough to allow for detection and action.
2. Scheduled Restoration (SR) / Scheduled Discard (SD) Tasks /
Preventive Maintenance (PM)
Definition: A scheduled task to overhaul a component (Restoration) or
replace it (Discard) at a fixed time or usage interval.
Objective: To prevent functional failure by renewing the item's "life
clock" before it reaches the end of its useful life.
Technical Applicability Criteria: An SR/SD task is applicable ONLY if:
1. The item has a identified and dominant age-related failure
pattern (Pattern A or B).
2. Most of the items survive to that age (a clear "wear-out" zone
exists).
3. The task restores the original inherent resistance to
failure without introducing a high risk of infant mortality.
3. Failure-Finding Tasks (FFT)
Definition: A scheduled task to check a hidden-function component to
determine if it has failed.
Objective: To reduce the risk of a multiple failure by discovering that a
protective device is in a failed state and restoring it to working order.
Technical Applicability Criteria: An FFT is applicable ONLY
to components with hidden functions (protective devices, standby
systems). It is not a preventive task.
o P-F Intervals, Useful Life, and MTBF
These three concepts are often confused but are fundamentally different.
Understanding them is critical to setting correct task frequencies and selecting
the right task type.
1. The P-F Interval (The Domain of Condition-Based Tasks)
What it is: The time interval between the point when a Potential
Failure (P) is first detectable and the point when it deteriorates into
a Functional Failure (F).
Visual: [Imagine a graph where condition deteriorates over time. The P
point is where the decline becomes detectable, and the F point is where it
crosses the failure threshold. The time between P and F is the P-F
Interval.]
How it's used: The task frequency for a CBT must be LESS than
the P-F Interval. If you inspect every 3 months, the P-F interval must be
reliably longer than 3 months to guarantee you find the problem before
failure.
o Example: If a bearing begins to spall (P) and will seize (F) after 400
hours of operation, your vibration analysis route must be performed
more frequently than every 400 hours (e.g., every 300 hours).
2. Useful Life (The Domain of Scheduled Restoration/Discard Tasks)
What it is: The typical operating age (time in service) at which an item
exhibits a rapid increase in the probability of failure due to age (wear-out).
This is the "knee" of the curve for Failure Patterns A and B.
How it's used: The task interval for an SR/SD task is set at a point
somewhat SHORTER than the useful life. This is to ensure the item is
replaced or overhauled before it enters the high-probability wear-out
zone.
o Example: If statistical analysis shows that a population of fan belts
has a useful life of 18,000 hours, the scheduled discard task might
be set at 15,000 hours.
3. Mean Time Between Failures (MTBF) (A Measure of Reliability, NOT a
Scheduling Tool)
What it is: A historical metric measuring the average operating time
between failures for a repairable item. MTBF = Total Operating Time /
Number of Failures.
The Critical Misconception: MTBF should NOT be used directly to
set a maintenance interval.
Why? MTBF is an average. It tells you nothing about the distribution of
failures. For the vast majority of components (showing random Failure
Patterns C, D, E, F), replacing a component at its MTBF is meaningless and
wasteful, as it has the same probability of failing immediately after
replacement as it did before.
o Example: If a light bulb has an MTBF of 10,000 hours, replacing all
bulbs every 10,000 hours would be incredibly wasteful, as many
bulbs will last much longer, and some will fail much sooner. The
failure pattern is random, not age-related.
Comparative Table:
Concep Task Type
Definition Primary Use in RCM
t Link
Time from
P-F To set the maximum
detectable Condition-
Interva allowable frequency for a
warning to Based Tasks
l Condition-Based Task.
functional failure.
To set the interval for a
Useful The age at which Scheduled
Scheduled Restoration/Discard
Life wear-out begins. Tasks
task.
To assess reliability, calculate
Not a direct
The average time availability, and inform
MTBF task-setting
between failures. the Likelihood in risk
tool.
assessment.
Practical Session: Selecting Tasks and Setting
Frequencies
Scenario: Critical Charge Pump (API Std 610)
Failure Mode: Mechanical Seal Failure leading to uncontrolled hydrocarbon
release (H₂S).
Consequence: Safety/Environmental. A proactive task is mandatory.
Option 1: Scheduled Discard Task
Proposal: Replace the seal every 24 months.
Evaluation:
o Is it Applicable? Only if seal failure is age-related. Analysis of
historical data shows no correlation between seal life and time;
failures are random due to flush fluid upsets. NOT APPLICABLE.
o Conclusion: Rejected.
Option 2: Condition-Based Task
Proposal: Use Airborne Ultrasound to listen for turbulent gas leakage
across the seal faces, which is an early sign of seal degradation.
Evaluation:
o Is it Applicable?
Can it detect reduced resistance? Yes. Turbulence indicates
face separation or damage.
Is the P-F interval consistent? Yes. From first sign of leakage
to catastrophic failure is typically 6-8 weeks in this service.
Is the P-F interval long enough? Yes. 6-8 weeks provides
ample time.
o APPLICABLE.
o Is it Effective? The cost of the program is minimal versus the infinite
cost of a safety incident. EFFECTIVE.
Setting the Frequency: The P-F interval is 6-8 weeks. To ensure
detection, the task frequency must be less than this.
Selected Task & Frequency: Perform airborne ultrasonic
inspection of the seal every 4 weeks.
Scenario: Emergency Diesel Generator
Failure Mode: Failure to start on demand due to a flat battery.
Function: Start automatically upon loss of power (a hidden function).
Consequence: Hidden -> Multiple Failure = Operational/Safety.
Task Selection:
A Condition-Based Task to "prevent" battery failure is not perfectly
applicable (batteries can fail suddenly).
A Scheduled Discard task is applicable (batteries have a useful life of ~5
years).
However, the primary strategy for a hidden function is a Failure-Finding
Task.
Selected Task & Frequency: Perform a simulated start (failure-
finding test) of the diesel generator monthly. This tests the entire
starting system, including the battery. The battery itself is also replaced
on a scheduled discard task every 5 years.
Practical Steps to Apply from This Module:
1. Calculate a Real P-F Interval: Review your CMMS and PdM data for a
component that failed. Can you find the first PdM alert (P) and the date of
failure (F)? Calculate the actual P-F interval. Use this to validate or adjust
your current route frequency.
2. Challenge an Age-Based PM: Find a time-based replacement task.
Graph the time-to-failure for the last 10 replacements. Does a clear wear-
out age exist, or are the failures scattered? This will prove if the task is
truly applicable.
3. MTBF vs. Useful Life Exercise: Take a population of 10 identical
pumps. Assume their failure times (in months) are: 5, 8, 12, 14, 16, 18,
22, 24, 28, 60.
o Calculate the MTBF: (5+8+12+14+16+18+22+24+28+60)/10
= 20.7 months.
o Now, look at the data. Is there a useful life? 8 of the 10 pumps failed
between 12-28 months. Setting a replacement at 18-20 months
might make sense. But replacing all pumps at the MTBF of 20.7
months would have been too late for half of them and far too early
for the one that lasted 60 months. This illustrates the difference.
Key Takeaways for 6.1:
Task selection is governed by strict technical criteria, not tradition.
You must match the task type to the failure pattern.
The P-F Interval is the fundamental concept behind condition
monitoring. Task frequency must be shorter than the P-F interval.
Scheduled tasks are only valid if a clear "useful life" can be
demonstrated. Most complex equipment does not have one.
MTBF is for measuring performance and assessing risk, NOT for
scheduling maintenance. Using it to set intervals is a classic and costly
error.
Module 6.2: RCM Logic and Decision Worksheets
Learning Objective: To master the use of the RCM logic tree for consistent
task selection and to understand the valid default actions when no proactive
task is applicable, ensuring all decisions are documented in a clear, auditable
worksheet.
o Task Selection Logic
The RCM logic tree is a standardized flowchart that forces the team to ask a
consistent series of yes/no questions about each failure mode. This ensures
discipline and eliminates personal bias. The logic flows directly from
the consequence of failure.
The RCM Logic Tree (Abridged for Clarity):
For each failure mode, starting with its consequence:
1. Is the failure evident to the operating crew?
o No (Hidden Failure) → The strategy is a Failure-Finding Task.
(Proceed to set the task interval).
o Yes (Evident Failure) → Proceed to question 2.
2. Does the failure have a Safety or Environmental consequence?
o Yes → You must find a proactive task. Ask: Is an age-based
(SR/SD) task technically applicable?
Yes → Implement the task. (Rare for safety consequences).
No → Ask: Is a condition-based (CBT) task technically
applicable?
Yes → Implement the task.
No → The only default action is REDESIGN.
o No (Operational or Non-Operational Consequence) → Proceed
to question 3.
3. Does the failure have an Operational consequence?
o Yes → A proactive task is desired if cost-effective. Ask the same
sequence as for safety: Is SR/SD applicable? If not, is CBT
applicable?
If a task is found and it's cost-effective → Implement it.
If no task is found or it's not cost-effective → Default actions
are Redesign or Run-to-Failure with a plan.
o No (Non-Operational Consequence) → Proceed to question 4.
4. Does the failure have a Non-Operational consequence?
o Yes → A proactive task is only justified by cost. Ask: Is a proactive
task cost-effective? (Is the cost of the task less than the cost of
the repair?)
Yes → Implement the least expensive applicable task.
No → The default action is RUN-TO-FAILURE.
o Default Strategies and "No Scheduled Maintenance" Scenarios
A "no scheduled maintenance" decision is not a failure of the RCM process; it is
a valid and often optimal outcome when properly justified by the logic. The
key is that it must be a deliberate decision, not an oversight.
1. Run-to-Failure (RTF) as a Default Strategy:
When is it Valid? RTF is the default and recommended strategy for
failures with Non-Operational Consequences where no proactive task
is cost-effective. It is also a possible default for Operational
Consequences if no proactive task exists and redesign is not justified.
What it DOES NOT mean: "We do nothing."
What it DOES mean: "We accept the failure but will manage its
consequences."
Required Supporting Actions (The "RTF Plan"):
o Ensure spare parts are available to minimize downtime.
o Have clear procedures for the repair.
o Train personnel so the repair can be done safely and efficiently.
o In some cases, it may mean having a redundant system already in
place.
2. Redesign as a Default Strategy:
When is it Mandatory? For Safety or Environmental
consequences where no proactive task is technically applicable.
When is it an Option? For Operational Consequences where the cost
of the failure is so high that redesign is a cheaper long-term solution.
Types of Redesign:
o Equipment Upgrade: e.g., replacing a single mechanical seal with
a tandem seal system (API RP 682).
o System Modification: e.g., adding a full equipment spare.
o Procedural Change: e.g., changing operating limits to avoid a
damaging process condition.
Practical Session: Completing an RCM Decision
Worksheet
The RCM worksheet is the formal record of the entire analysis. It provides
traceability and is essential for auditing and future updates.
Sample RCM Worksheet (Abridged)
Functi Ta
Item / Proacti
onal Failure Effects & sk Default
Functio ve
Failur Mode Consequence Ty Action
n Task
e pe
Vibrat
Charge
ion
Pump Effects: Unit
Analy
P-101A Bearing trip,
sis
Functio seizure production
(mont (Not used
n: No due to loss CB
hly) - task
Transfe Flow lube oil ($150k/day). T
Oil selected)
r 800 contami Consequenc
Analy
GPM @ nation e:
sis
450 Operational
(quart
psig
erly)
Functi Ta
Item / Proacti
onal Failure Effects & sk Default
Functio ve
Failur Mode Consequence Ty Action
n Task
e pe
PSV on Effects: Hidd
V-101 en. Multiple
Fails Bench
Functio failure leads
to Internal test
n: Open to vessel (Not used
open corrosio every FF
at 250 rupture - task
on n/ 24 T
psig to (BLEVE). selected)
dema fouling mont
prevent Consequenc
nd hs
overpre e: Hidden ->
ssure Safety
Standb
Effects: Redu
y
ndancy lost. If (No
Cooler
main fan fails, cost- Run-to-
Fan
temp rises, effecti Failure (B
Functio Fails
Motor minor ve N/ ut ensure
n: Start to
burnout production proacti A spare
automa start
impact. ve motor is in
tically
Consequenc task stock)
when
e: Hidden -> found)
temp >
Operational
90°C
Pipelin Ruptu External Effects: Majo (No N/ REDESIG
e re corrosio r spill, fire, proacti A N (e.g.,
Section n environmental ve Implement
Functio causing damage. task Cathodic
n: wall Consequenc can Protection
Contain thinning e: reliabl system &
crude Safety/Envir y improve
oil at onmental detect coating)
1200 all
psig extern
Functi Ta
Item / Proacti
onal Failure Effects & sk Default
Functio ve
Failur Mode Consequence Ty Action
n Task
e pe
al
corrosi
on on
buried
pipe)
Workshop Exercise: Analyze a Pressure Transmitter
Asset: Pressure Transmitter PT-101 on a Fuel Gas line.
Function: Provide a 4-20 mA signal to the control system representing
pressure (0-500 psig).
Functional Failure: Provides an inaccurate signal (reads 100 psig when
actual is 400 psig).
Failure Mode: Sensor drift due to process contamination.
Failure Effects: Control system receives wrong data, potentially leading
to incorrect valve positioning.
o End Effect: Inefficient combustion, or in a worst-case, a unit trip
due to incorrect pressure control. Consequence: Operational.
Apply the Logic Tree:
1. Evident? No, the failure is hidden from the operator; it requires a test to
discover. The signal still appears valid.
2. Hidden Function? Yes. It is a monitoring device whose failure is not
evident.
3. Logic Path: For a hidden function, the strategy is a Failure-Finding
Task.
Worksheet Entry:
Proactive Task: "Calibrate transmitter PT-101 every 12 months." (This is
a Failure-Finding Task to discover if it has "failed" out of its tolerance).
Task Type: FFT
Default Action: N/A
Practical Steps to Apply from This Module:
1. Logic Tree Walk-Through: As a team, take a blank logic tree and a
known failure mode. Verbally walk through every question, justifying each
"yes" or "no" answer. This builds deep familiarity with the process.
2. Worksheet "Blitz": Take a system with 5-10 known failure modes. In a
focused 2-hour session, work as a team to complete just the "Proactive
Task" and "Default Action" columns for each one, using the logic tree as
your guide. This builds speed and consistency.
3. The "Redesign" Brainstorm: Identify one failure mode in your plant
with safety consequences. Assume that no proactive task is possible. Hold
a brainstorming session: "If we must redesign this to make it safe, what
are our options?" This trains the muscle for the most critical default
action.
Key Takeaways for 6.2:
The RCM logic tree is a disciplined, consequence-driven
checklist that ensures consistent and defensible maintenance decisions.
"No Scheduled Maintenance" (Run-to-Failure) is a valid, logical
output for non-operational consequences, but it requires a plan to
manage the failure.
Redesign is a mandatory default action for safety-related
failures when no proactive task can be found. It is not a failure but a
necessary engineering response.
The RCM worksheet is the legal document of the analysis, providing
a clear audit trail from function to task, which is crucial
for PSM compliance and knowledge retention.
Module 6.3: Practical Session - Case Study: FMEA and
Maintenance Strategy Selection
Learning Objective: To apply the complete end-to-end RCM process, from
Functional Analysis and FMEA to maintenance task selection, using a realistic
case study. This session consolidates all learning from previous modules into a
single, hands-on exercise.
Case Study: Seawater Injection Pump System
Background:
You are part of a reliability team at an offshore platform. The Seawater Injection
System is critical for reservoir pressure maintenance. A failure can lead to a
significant reduction in oil production within 48 hours. The system uses high-
duty centrifugal pumps. We will focus on one pump, SW-Pump-101A, which is
critical with no immediate spare.
Relevant Standards & Data:
Pump Design: API Std 610 (Centrifugal Pumps)
Seals: API RP 682 (Shaft Sealing Systems)
Corrosion: NACE/AMPP standards for seawater service.
Historical Data from CMMS: MTBF for this pump type is 36 months. The
dominant failure modes are bearing and seal related.
Step 1: Functional Analysis (RCM Questions 1 & 2)
Instructions: Define the primary and secondary functions, then identify the
functional failures.
Function & Performance
Function Type Functional Failure
Standard
Inject filtered seawater at a Complete: Fails to inject
rate of 1200 m³/hr against (0 m³/hr).
Primary Function
a discharge pressure of 85 Partial: Injection rate
bar. falls below 1000 m³/hr.
Contain seawater within
Secondary Function External leakage of
the pressure boundary with
(Containment) seawater.
zero external leakage.
Interface with control
Secondary Function system (provide 4-20mA Provides inaccurate or
(Control/Safety) signal for discharge zero signal to DCS.
pressure).
Secondary Function Operate with a pump Pump efficiency drops
(Efficiency) efficiency of >78%. below 75%.
Step 2: Failure Modes, Effects, and Consequence
Analysis (RCM Questions 3, 4, & 5)
Instructions: For the functional failure "Fails to Inject (0 m³/hr)", we will
conduct a focused FMEA. Complete the following table as a team.
FMEA Worksheet (Abridged)
Item Function Functional Failure Failure Mode (Root Failure Effects (Local -> System -> End Severit Occurrence Detectio RP Consequenc
Cause) Effect) y (1- (1-10) n (1-10) N e Category
10)
SW-Pump- Inject 1200 m³/hr @ 85 Fails to Inject (0 1. Coupling Local: Coupling breaks, loud noise. 8 3 3 72 Operational
101A bar m³/hr) Failure due to System: Pump stops, discharge
severe pressure drops to zero, high vibration
misalignment. alarm.
End Effect: Injection stops. Production
decline within 48 hours. Cost:
~$500k/day.
2. Bearing Local: Bearing overheats and locks up. 8 4 4 128 Operational
Seizure due to System: Pump shaft stops rotating,
lubricant motor trips on overload.
contamination from End Effect: Injection stops. Production
water ingress. decline. Cost: ~$500k/day + $50k
repair.
3. Pump Local: Impeller pitting and damage, 7 5 3 105 Operational
Cavitation due to noise.
clogged suction filter System: Flow becomes erratic,
(from marine pressure drops, pump may trip.
growth). End Effect: Reduced injection rate
(<1000 m³/hr) or complete stop.
Production impact. Cost: ~$250k/day.
Severity Scale: 1 (None) to 10 (Catastrophic: Multiple fatalities, total loss)
Occurrence Scale: 1 (Very Low: <0.001/yr) to 10 (Very High: >1/yr)
Detection Scale: 1 (Very High: Certain to detect) to 10 (Very Low: Cannot detect)
Step 3: Maintenance Task Selection (RCM Question 6 & 7)
Instructions: Using the RCM logic tree and the information above, select an appropriate maintenance strategy for each failure mode. Justify your choice based on
applicability and effectiveness.
Failure Mode Consequen Applicable & Effective Proactive Task Selected Strategy & Justification Default Action (if no
ce task)
1. Coupling Operational Condition-Based: Laser shaft alignment check during Strategy: Laser alignment check quarterly. Run-to-Failure with a
Failure monthly maintenance. Justification: Applicable - Misalignment is detectable before failure. spare coupling on
Condition-Based: Visual inspection for signs of wear. Effective - Cost of alignment is minimal vs. production loss. site.
2. Bearing Operational Condition-Based: Vibration analysis to detect early Strategy: Monthly vibration analysis + quarterly oil analysis. (Not Applied)
Seizure bearing wear. Justification: SR/SD not applicable (failure is random, not age-
Condition-Based: Oil analysis to detect water based). CBT is highly applicable and cost-effective given the high
contamination. consequence.
Scheduled: Time-based bearing replacement.
3. Pump Operational Condition-Based: Ultrasonic flow meter to detect Strategy: Monitor differential pressure (dP) across suction (Not Applied)
Cavitation erratic flow / differential pressure monitoring across filter daily. Clean filter when dP exceeds 1.5 bar.
suction filter. Justification: A scheduled task (cleaning) is applicable and effective
Preventive: Scheduled cleaning of suction filter. here, as fouling is time/usage-based. Condition monitoring (dP)
provides the trigger.
Facilitated Discussion & Key Insights
1. Linking FMEA to RPN and Consequence:
o Notice that Bearing Seizure has the highest RPN (128), which
correctly flags it as a high-priority issue. The RCM process takes this
a step further by using the Operational Consequence to mandate
a proactive task.
2. Justifying Condition-Based over Time-Based:
o For the bearing, a common traditional approach might be "overhaul
pump every 3 years." The RCM analysis shows this is not applicable
(failure is random) and not the most effective strategy. Condition-
based tasks are superior.
3. The Power of a Simple Monitoring Parameter:
oFor cavitation, the team selected a simple, reliable, and direct
parameter—Differential Pressure (dP)—rather than a more
complex or expensive technology. This is a hallmark of a well-
thought-out RCM strategy.
4. Documentation for Audit and Compliance:
o The completed worksheet now serves as a defensible record. It
shows regulators (e.g., for OSHA PSM) that a systematic, risk-
based approach was used to determine the maintenance strategy
for this safety-critical system.
Practical Steps to Apply from This Module:
1. Run Your Own Mini-Case Study: In your team, select a single, well-
understood piece of equipment (e.g., a small pump, a compressor
auxiliary system). Replicate this 3-step process in a 2-hour workshop.
o Step 1 (15 min): Define Functions and Functional Failures.
o Step 2 (45 min): Brainstorm Failure Modes and Effects for one
functional failure.
o Step 3 (60 min): Use the logic tree to select tasks.
2. The "Why" Challenge: For each selected task, have a team member
play "devil's advocate" and ask "Why is this the best task?" The facilitator
must defend the choice using the concepts of applicability (P-F interval,
age-relatedness) and effectiveness (cost-benefit).
3. Management Briefing: Use the completed worksheet to create a 5-slide
presentation for management, justifying the potential investment in new
condition monitoring technologies (e.g., a vibration analyzer) based on the
avoided operational losses documented in the FMEA.
Key Takeaways for 6.3:
FMEA is the "engine" that identifies and prioritizes failure
modes, while the RCM logic tree is the "driver" that selects the
correct strategy.
The completed RCM/FMEA worksheet is a powerful communication
and compliance tool, providing a clear, logical audit trail from function
to task.
The most effective strategy is often a combination of task
types (e.g., condition monitoring to trigger a scheduled restoration), as
seen in the cavitation example.
Practicing this end-to-end process on a real asset is the best way
to cement the RCM methodology and build confidence to tackle more
complex systems.