0% found this document useful (0 votes)
19 views26 pages

Problem Management Process Overview

The Problem Management Process outlines procedures for identifying, logging, and resolving IT service problems to prevent future incidents. It includes policies on problem logging, prioritization, ownership, and escalation, as well as roles and responsibilities for involved parties. The process aims to improve service quality, minimize incidents, and ensure effective communication throughout the problem lifecycle.

Uploaded by

kawumin
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
19 views26 pages

Problem Management Process Overview

The Problem Management Process outlines procedures for identifying, logging, and resolving IT service problems to prevent future incidents. It includes policies on problem logging, prioritization, ownership, and escalation, as well as roles and responsibilities for involved parties. The process aims to improve service quality, minimize incidents, and ensure effective communication throughout the problem lifecycle.

Uploaded by

kawumin
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd

Problem Mnagement Process

Table of Contents

Problem Mnagement Process.............................................................................1


Table of Contents................................................................................................2
Document Control................................................................................................4
Document Statistics................................................................................................ 4
Revision History...................................................................................................... 5
Process Overview.................................................................................................................... 6
1. Description and Scope....................................................................................6
1.1 Scope:............................................................................................................... 6
1.2. Objective.......................................................................................................... 7
2. Policies................................................................................................................................ 8
Log a Problem.....................................................................................................8
Delayed Logging of a Problem............................................................................8
Priority.................................................................................................................8
Altering the Priority..............................................................................................8
Ownership...........................................................................................................9
Parallel Activities.................................................................................................9
Reject a Problem.................................................................................................9
Notification...........................................................................................................9
Escalation..........................................................................................................10
Root Cause........................................................................................................11
Unresolved Problems........................................................................................11
Known Error.......................................................................................................11
Close a Record..................................................................................................12
Monitoring Progress..........................................................................................12
Interfacing Processes........................................................................................12
3. Roles and Responsibilities................................................................................................ 13
4. Key Performance Indicators (as per contract or mutually agreed target)...........................15
5. Workflow............................................................................................................................ 16
Detect, Log, Categorise & Prioritise - Problem Activity Flow.............................16
Detect, Log, Categorise & Prioritise – Problem Activity Description.................17
Investigate and Diagnose – Problem Activity Flow............................................20
Investigate and Diagnose - Problem Activity Description..................................21
Resolve and Close - Problem Activity Flow.......................................................23
Resolve and Close - Problem Activity Description............................................24
Appendix A: RCA Template...............................................................................26
Document Control
Document Statistics
Type of Information Document Data
Title Problem Management
Document Revision #1.0 1.0
Last Date Document was Updated
Total Number of Pages
Document Filename
Document Creator

Document Approver

Document Reviewer(s)

Document Distribution List


Revision History

S. Version Revision Date Date


Nature of Change Approved Released
No. No. Date
Process Overview

1. Description and Scope


The purpose of the Problem Management process is to resolve problems affecting the IT
service, both reactively and proactively. Problem Management finds trends in incidents,
groups those incidents into problems, identifies the root cause of problems, and tracks
change requests (RFCs) against those problems.
Problem Management manages the resolution and prevention of problems that affect the
normal operation of a company’s information technology (IT) services. This includes
ensuring that all failures are corrected, preventing any recurrence of those failures, and
using preventative measures to reduce the number of failures.
Definition of problem: “Unknown cause of one or more incidents” (Source: ITIL® V3 Glossary).
Note: The cause is not usually known at the time a problem record is created, and the
Problem Management process is responsible for further investigation. Once root cause of
the problem is known and a workaround or a permanent solution has been identified and
provided, the problem record can be resolved. Each problem record documents the
details and lifecycle of a single problem
1.1 Scope:
Entry into the Problem Management process is triggered where a problem is detected
pro-actively or reactively.
Problems are proactively detected as a result of scheduled or ad hoc analysis of
collected data.
Problems are reactively detected where, during the analysis of incident data, it is
identified that further investigation is required into the root cause of one or more
incidents.
This detection triggers the logging of a problem in the IT Service Management tool.
Notes:
Proactive function is to identify and solve problems before incidents occur. This
function involves identifying problems and resolving them before they lead to any
incidents that disrupt normal operations.
Reactive function is to solve problems relating to one or more incidents. This involves
resolving problems that are the root causes of incidents and preventing further incidents
based on the problem.

The following will trigger an exit from the Problem Management process:
 Appropriate actions taken to resolve a problem to prevent future incidents and
problems and,
 The activity of closing the problem record and a Major Problem Review initiated
where criteria met.
The scope of the Problem Management process includes a series of activities which
involve the following:
 Categorisation and prioritisation of problems
 Investigation (root cause analysis) into underlying cause and identification
 Progressing problems into known errors where underlying root cause identified and
workaround available
 Solution (and workaround) definition and selection
 Where justified, resolving known errors by applying a permanent solution to prevent
any recurrence
 Appropriate prioritisation of resources required for resolution based on business
need
 Submission of change requests to implement solutions
 Continuous monitoring and tracking of problems and known errors (taking
appropriate action where required)
 Communicating problem information.
 Contribution to the collective problem resolution knowledge base and solution
creation
 Formulating recommendations for improvement, maintenance of problem records
and review of the corrective actions
 Updating records with problem and known error information (KEDB).

The scope of Problem Management excludes the following:


 Identification, creation and resolution of incidents (Incident Management)
 Actual implementation of the resolution of a problem as Problem Management
initiates the resolution through Change Management where criteria met
 Knowledge Management methodology. Knowledge Management
Methodology implies to creation and maintenance of Knowledge
articles. Known Error Database (KEDB) is a subset to the overall
Knowledgebase. Problem Management will focus on KEDB.
1.2. Objective
The objectives of the Problem Management process are:
 Identification and resolution of problems proactively or reactively
 Potential incidents or the recurrence of incidents are prevented
 Number and adverse impact of incidents and problems is minimized
 Where possible, root cause of a problem and its subsequent resolution is established
 Reduction of incidents and problems to an acceptable risk at an acceptable cost
 Duration of problem life cycles are reduced
 Improved communication of problems with all involved parties
 Improved customer service and service quality
 Maximized system availability and improved service levels
2. Policies
Log a Problem
Problem Tickets need to be logged for every Sev1 and Sev2 Incident and displayed as related tickets,
having a correlation between them. All relevant detail of RCA like symptoms, Cause and resolution
actions need to be there along with preventive action detail.

It is vital that all information entered in to the IT Service Management tool is accurate and follows the
guidelines set in this document.
Upon the detection of a problem either proactively or reactively, Problem Management process must
be invoked immediately to log the problem with the following minimum information:
 Unique Problem record identifier
 Date and Time Problem recorded
 Description of problem and any actions already taken to assist with investigation
 Priority as per policy
 Configuration item(s) involved
 Classification (category)
 Indication, where problem has been proactively identified or is a recurrence.
 Any Third Party record number for referencing purposes
Where a problem is reactively detected, a relationship must be built by linking the logged
problem to all associated incidents.
To log a problem in the IT Service Management tool, one must have the appropriate tool access.

Procedure to log a Problem Ticket is explained in the attachment as Appendix B on Page 30.

Delayed Logging of a Problem


If the logging of a problem is delayed due to the IT Service Management tool being
unavailable, then the problem is logged manually either on paper or computer for audit
tracking purposes.
Verbal notification must be performed where automated notification is also unavailable.
When IT Service Management tool becomes available, all information must then be
logged in a problem record, including the backdating of any date and times to ensure
integrity of data, as soon as possible and no later than 2 business days.
Refer Appendix A IT SMT Manual Logging Template

Priority

Each problem must have an allocated priority level.


Where multiple incidents exist, the priority of all these incidents combined must be taken
into consideration when determining the correct problem priority.

The priority of the problem is inherited from the associated incident. Ref priority Policy in
incident management process for the priority of incident

Altering the Priority


Any reassessment of the problem priority must occur prior to the resolution of a problem.
Prior approval must be obtained from the Problem Manager before altering the priority of
a problem.
For all problems, where the problem priority has been altered, the reason and the
obtained approval must be entered in the problem record.

Ownership
Ownership is established by assigning a record to the appropriate support group regardless of formal
acceptance or acknowledgment.
There is only one owner of a record at any given time, and that is the support group it is currently
assigned to.
Any reassignment of a record may occur throughout the Problem Management process.
Prior to any record reassignment the owning support group ensures that the reason and all actions
leading up to the reassignment have been documented in the record.
Perform verbal notification upon record assignment to ensure awareness and no delays are
experienced.
Where ownership disputes cannot be resolved, immediate escalation to the Problem Manager must
occur.

Parallel Activities
Other Support Groups may be engaged to assist and perform parallel activities where record
ownership resides with another support group.
Irrespective of record ownership, where parallel activities are required, a problem is to be treated
based on its allocated priority.
Where a Support Group is engaged in parallel activities, it is their responsibility to document all
investigation activities and findings in the problem record.
Where a Third Party is engaged in parallel activities and does not have direct IT Service Management
tool access, then the responsibility for ensuring updates to the problem record remains with a support
group that owns the relationship (i.e. acting on behalf of the Third Party).
Where a request for parallel activities is refused, immediate escalation to the Problem Manager must
occur.

Reject a Problem
When ownership is rejected by the assigned Support Group, the record must be
updated to indicate the reject reason and then either:
Reassign the problem to the correct Support Group (where known) or
 Contact Problem Manager to establish an appropriate owner
Verbal notification must accompany any reassignment to ensure no further delays are
experienced.
Where ownership disputes occur, escalate to the Problem Manager.

Notification
Notification occurs throughout this Problem Management process. This ensures that the
responsible role is notified about further action required by them, or to communicate
relevant information on a problem’s progress.
All notifications performed including any delays experienced must be recorded in the
problem record.
The table below outlines the key notifications by the various roles involved in the Problem
Management process.
The ClientIT and any other Service Providers upon receiving notification from Vendor are responsible
for communicating internally within own organization.
If the The Notifies the By
Analysis highlights a service Support Group Problem Manager E-mail communication to
improvement opportunity central mailbox Vendor’s
Problem Mngt Team
Priority of a Problem requires Support Group Problem Manager Verbal notification to obtain
alteration prior approval
If the The Notifies the By
Record ownership is Rejected Existing Support Support Group to Verbal notification upon
and correct ownership is Group which the record is rejecting ownership
known being reassigned Record updated with reject
reason
Existing Support Problem Manager Verbal notification upon
Group rejecting ownership
Record updated with reject
reason
Workaround exists for a Support Group Problem Manager Record updated
problem E-mail communication
Problem Manager Stakeholders Verbal notification and
email notification
Known Error solution requires Support Group Other Support Group Verbal and email
involvement from other and managed 3rd communication
analysts or third parties Party Suppliers Record updated
Proposed solution requires Support Group Account Team and E-mail notification and
consultation with ClientIT Problem Manager Record updated
Account Team ClientIT Verbal notification

Figure 1 Problem Notification

Escalation
Escalation occurs throughout the Problem Management process. This ensures additional
attention and visibility is given to issues so as to meet service level agreements and or
customer expectations.
Escalation may be functional or hierarchic and the level of escalation is dependent on the
nature and significance of the issue.
All escalations must continue until an agreed outcome is achieved. The outcome of the
escalation must be documented in the problem record and communicated accordingly.
The table below outlines the key escalations by the various process roles involved in the
Problem Management process.

If the escalation is related to Then the Initially escalates to Escalation Mechanism


Parallel activities or investigations Support Group Other Support Group  Verbal escalation
by other support groups  Problem record
updated
The underlying cause of a Support Group Problem Manager E-mail escalation to
problem remaining unidentified central mailbox Vendor’s
Problem Mngt Team to
determine the course of
action
Ownership disputes, priority All Process Roles Problem Manager Verbal escalation
issues or the correct owning group Record updated
cannot be determined
Problem Manager Account Team Verbal escalation
Record updated
An identified entitlement issue All Process Roles Account Team Verbal escalation and
follow up e-mail

Progression of a problem or Problem Manager  Line Management Verbal escalation


Known Error or any risk to pre-  Account Team Record updated
defined targets
A problem not progressing and Problem Manager Support Group  Verbal notification /
requires actions to be performed e-mail
If the escalation is related to Then the Initially escalates to Escalation Mechanism
Justification provided for not Support Group Problem Manager E-mail notification
performing root cause analysis or
providing permanent solution
Client dissatisfaction All Process Roles Account Team Verbal escalation

Table 6 Problem Escalation

Root Cause
A support group must conduct a Root Cause Analysis (RCA) for each problem
regardless of priority.
Where the cause of a problem cannot be identified, the assigned Support Group is to
establish preventative measures so as to capture additional data to assist with root
cause identification or prevent future incidents.
RCA need to have its scope ( In or out of Lycamoney) defined specifying severity (S1or
S2) as well
The Problem Manager will review and accept the results prior to communicating the
information to relevant stakeholders.
RCA for the Priority 2 problem record must be filled in the tool only, having Symptom,
Cause, Resolution details and all the identified corrective and preventive actions
(CAPA). The identified actions items must be tracked in the ITSM tool.
An exemption from both the Problem Manager and Account team must be obtained
where RCA will not be performed, as it’s deemed as not warranted or justified. All
relevant information (i.e. the reason, exemption granted or not) must be documented in
the problem record.

Unresolved Problems
A problem remains unresolved where:
 The root cause of a problem was investigated but the underlying cause was not
identified.
 Analysis into the root cause of the problem was not warranted or justified based on
the impact and urgency assessment and incident count(s).
 Root cause is identified and solution proposed, however an agreement has been
reached between the Account Team and ClientIT not to resolve the known error.

Where a problem remains unresolved a record must not be closed so as to ensure any future
incidents are identified, matched and tracked against the associated problem.
Unresolved problem records that have remained open for a defined period of time (i.e. 13 months) will
be reassessed by the Problem Manager
The Problem Manager must document the assessment outcome in the problem record
and close the record only after agreement reached post discussion with Account Team.

Known Error
A problem progresses to a known error once a root cause and workaround has been successfully
diagnosed and documented.
To resolve and eliminate a known error, an appropriate solution must be identified. This
solution is evaluated as to whether eliminating the known error is feasible and justifiable.
Prior to any implementation, a solution must be reviewed and approved by both technical
and operation stakeholders. The approved solution activities are formally documented
and implemented in accordance with the Change Management process (where criteria
met).
Where the solution is not to resolve and eliminate a known error, document the reason
and the agreement obtained within the account and operations team
Close a Record
Problem Manager must review and approve the resolution of all problem records. All approvals must
be documented in the record prior to closure.
Once a record has been resolved any recurrence of the same or similar problem must be recorded as
a brand new problem and a relationship established.
A problem record may be cancelled only where it is was opened in error and no problem activities
have commenced. Prior to cancellation provide cancellation reason in the record.

Monitoring Progress
Problem Manager must actively monitor the progress of all problems and known errors to
ensure their progression within lifecycle of the Problem Management process.
Specific attention must focus on those problems and known errors having an extended
period of inactivity or slow progression. Liaise with relevant party (i.e. Support Group,
Account Team) and initiate any required actions so as to progress the problem or known
error.
Problem Manager must ensure that the problem record is regularly and accurately
updated throughout its lifecycle.

Interfacing Processes
Problem Management process interfaces with the following processes:
1. Incident Management provides incident data for analysis so as to identify trends,
and to group related incidents that will assist in identifying whether a problem
already exists or requires creation. Problem Management provides information
relating to unresolved problems, Known Errors, available workarounds and
problem closure.
2. Problem Management will initiate a Request for Change (RFC) where it is
identified that a configuration item is to be altered in relation to a workaround,
permanent solution or preventative measure. Change Management provides the
status and the final outcome (i.e. successful or unsuccessful) of the implemented
change.
3. Run and Monitor Operation (Event Management) provides Problem
Management with real time and historical event information to assist with the
proactive detection of problem(s), the investigation of a problem and service
quality improvement.
4. Configuration Management provides Problem Management with configuration
item information to assist with problem investigation, resolution and preventative
measures. Problem Management provides information and modifications to
existing configuration items to enable the configuration data to be validated and
updated.
5. Key availability data is provided to Problem Management so that investigation,
resolution and preventative actions can be initiated to enable availability targets
to be met. Problem Management provides SLA Management Process with
problem information to allow for availability trends to be identified and remedial
action to be instigated.
Capacity information is provided to assist with problem investigation, resolution and pro-
active problem management activities. Problem information is provided to Capacity
Management to assist with capacity planning
3. Roles and Responsibilities
Responsibilities have been defined in relation to the Problem Management process activities using a
RACI model. In the table below are the defined RACI authority values.

R Responsible Individual(s) that are responsible for performing and completing a specific
activity.
The degree of responsibility is defined by the Accountable person.
(Note: Responsibilities may be shared)
A Accountable An individual who has the prime lead and the ownership and is ultimately
accountable for ensuring a specific activity is completed.
C Consulted Individual(s) who are consulted prior to a final decision or action.
(Two-way communication)
I Informed Individual(s) that receive information after a decision or action is taken. (One-
way communication)
Table 1 RACI Model
In the Roles and Responsibility matrix below, each responsibility has been defined
against functional role(s) with a defined RACI authority value.

Problem Manager
No. Responsibility Description

Support Group

Account Team
(Vendor)

(Vendor)

(Vendor)

ClientIT
Responsibilities for Detect, Log, Categorise and Prioritise
1 Re-actively detects a problem based on analysis of incident data R R,A
2 Proactively detects a problem based on analysis of information from R A
various other sources
3 Identify incidents not matched to problems or known errors R A
4 Identify and provide input into any service improvement opportunities R R R,A
5 Where a problem does not already exist, log a problem in the IT Service R R,A
Management tool with required information
6 Build relationships (link) the problem to associated incident(s) R R,A
7 Classify problem to determine category R R,A
8 Allocate the priority based on the impact and urgency assessment taking R R,A
into consideration the severity (seriousness) of problem
9 Identify and assign problem to the appropriate support group R R,A
10 Escalate where ownership cannot be determined or disputes exist R R,C A,C,
I
11 Assess whether investigation into the root cause of a problem is R R A
warranted or justified
12 Document the reason and communicate where root cause analysis is not R R,A C,I C,I
performed and the problem remains unresolved
Responsibilities for Investigate and Diagnose
13 Confirm ownership or reroute record where required R A
14 Gather additional information to commence investigation R,A
15 Obtain prior approval when altering the Priority of problem R A C C
16 Search for an available workaround, or create a workaround where one R,A
does not exist
17 Communicate workaround information R A,I I I
18 Update knowledge base with problem information R,A
19 Perform Root Cause Analysis to determine cause, engaging others in R R,A
parallel investigation
20 Notify where service levels are at risk R R,A I I
21 Formally document the findings of the RCA investigation R,A
22 Determine course of action where root cause is unidentified R A,C C
23 Convert problem into known error where cause is identified R A
Problem Manager
No. Responsibility Description

Support Group

Account Team
(Vendor)

(Vendor)

(Vendor)

ClientIT
24 Provide Incident Management (and Service Desk) with problem and R R,A
known error information
Responsibilities for Resolve and Close Problem
25 Diagnose known error and identify appropriate solution R,A
26 Develop a resolution plan outlining the key activities based on the R,A
proposed solution
27 Verify proposed solution with key stakeholders to determine the solution’s R A C C
feasibility and whether sufficient justification exists to resolve problem
28 Raise a proactive problem where the solution applies to the rest of the R,A
environment so as to prevent future incidents
29 Document reason where proposed solution is not feasible or justifiable R,A A C
and problem remains unresolved, handling any entitlement failures
30 Initiate Request for Change (RFC) or submit project proposals to R,A
implement agreed solution
31 Apply permanent solution as per the documented resolution activities R,A
32 Confirm elimination of Known Error by testing the applied resolution R,A
33 Complete problem record by validating that all key information was R,A
documented, adjust problem category where required
34 Update knowledge base with problem or known error information R,A
35 Close problem record as per policy and any other related problem records R,A I I
and perform communication of closure
36 Notify Incident Management to ensure all related incidents are also closed R,A
37 Initiate an internal Major Problem Review as per policy R,A
38 Obtain any required approval prior to closure of an unresolved problem R,A C C
39 Perform notification where Known Error remains unresolved R,A I I
Responsibilities across Problem Management activities
40 Establish and produce reports to communicate problem information R R,A
41 Provide regular progress updates in SMT record through to closure R A
42 Attend problem review meetings where required R R,A R
43 Facilitate customer meeting and distribute customised report as an out R,A I
from any problem review meetings
44 Manage record ownership disputes and issues R A,C
45 Continuously monitor and track the progress of problems and Known R,A
Errors through to closure
46 Identify problem records not progressing and initiate required actions , C R,A
communicating where required
Table 2 RACI Matrix
4. Key Performance Indicators (as per contract or mutually agreed
target)
[Link]. KPI description Measurement
RCA document Submission within 5 business days
75% *
1 timelines – for Severity 1 problem ticket

RCA sheet Submission within 7 business days timelines –


for Severity 2 problem ticket (Target to be baseline post 3 75% *
2 months)

RCA Action Item closure within agreed timelines in RCA


75%
3 document

Reduction in overall number of high severity incidents


25 % Half yearly (To be
including Recurring (S1/S2) – (baseline sev1 – 18 & sev2 –
reviewed post every 6 months)
4 27) Scope IT and post inclusions.

Aging analysis of the problems to be published monthly


Monthly 100%
5 (Open state)

10% every Quarter( measured


Enhance the KEDB (Known error database ) over a previous quarter
6 numbers)
* Business day starting next day of problem ticket raised
* Excluded for the Airtel’s Vendor dependent RCA’s
5. Workflow
Detect, Log, Categorise & Prioritise - Problem Activity Flow

Figure 2 Detect, Log, Categorise & Prioritise: Problem Activity Flow


Detect, Log, Categorise & Prioritise – Problem Activity Description

Step Activity Role Activity Description


1 Support Group Detect problems proactively by:
Reviewing or analysing information retrieved from various
sources (i.e. event, operational monitoring data, and supplier
information) so as to prevent potential incidents from
occurring.
Detect problems reactively by:
Analysing incident data received from Incident Management
process to identify incident trends or incidents that are
currently not matched (linked) to an existing problem or known
error.
During the detection of a problem, identify any opportunities
for service improvement and notify Problem Manager
Problem Manager 1. Detect a problem proactively or reactively by
reviewing data to identify:
2. Incident recurrences,
3. Incidents that are unmatched to existing problem(s)
or known error(s)
4. Any suspect configuration items of an IT
infrastructure.
5. During the detection of a problem, identify and or
receive any opportunities for service improvement.
6. Provide Account team with any proposals to improve
an IT service.
2 Account Team Receive any proposals relating to opportunities for service
improvement in the provision of IT services.
Manage the service improvement and invoke the relevant
process as required.
Note: Service improvement does not necessarily identify an
actual problem with an IT service, but outlines an opportunity
to improve it.
3 Support Group Where a problem is detected either proactively or reactively;
perform a search to determine whether a problem record
already exists or requires creation.
Problem record exists?
If Yes - Go to step 5 Link Problem to Associated Incidents
If No - Go to step 4 Log a Problem Record
Problem Manager Where a problem is detected either proactively or reactively;
perform a search to determine whether a problem record
already exists or requires creation.
Problem record exists?
If Yes - Go to step 5 Link Problem to Associated Incidents
If No - Go to step 4 Log Problem Record

4 Support Group Log the detected problem in the IT Service Management tool
with the required information.
Refer Policy: Log a Problem
Problem Manager Log the detected problem in the IT Service Management tool
with the required information.
Refer Policy: Log a Problem
5 Support Group Build a relationship by linking matched incidents to an existing
problem or newly logged problem record.
Problem Manager Build a relationship by linking matched incidents to an existing
problem or newly logged problem record
6 Support Group Based on the available information, classify the problem into
its correct category (i.e. Hardware, Software, Application)
Note: A problem may be re-categorised at any stage
throughout the Problem Management process.
Problem Manager Based on the available information, classify the problem into
Step Activity Role Activity Description
its correct category (i.e. Hardware, Software, Application)
Note: A problem may be re-categorised at any stage
throughout the Problem Management process.
7 Support Group Determine the impact of the problem by performing an
assessment as to what effect a problem has or may have on
the business and any service levels.
Determine the urgency of the problem by assessing how long
the business can tolerate the unresolved problem.
Take into consideration the severity and seriousness of the
problem from an IT infrastructure perspective.
As required, liaise with Problem Manager and or Account
Team to determine the correct impact and urgency of a
problem.
Refer Policy: Priority
Problem Manager Determine the impact of the problem by performing an
assessment as to what effect a problem has or may have on
the business and any service levels.
Determine the urgency of the problem by assessing how long
the business can tolerate the unresolved problem.
Take into consideration the severity and seriousness of the
problem from an IT infrastructure perspective.
As required, liaise with Account Team and Support Group to
determine the correct impact and urgency of a problem.
Refer Policy: Priority
Account Team As required, liaise with Support Group, Problem Manager and
ClientIT to determine the correct impact and urgency of a
problem.
ClientIT As required, confirm business impact and urgency in
consultation with Account team.

8 Support Group Allocate a Priority to a problem based on the combination of


the assessed business impact, its severity and the urgency.
Refer Policy: Priority
Problem Manager Allocate a Priority to a problem based on the combination of
the assessed business impact, its severity and the urgency.
Refer Policy: Priority
9 Support Group Ensure record ownership is assigned to relevant support
group.
Perform notification when reassigning record ownership.
Escalate to Problem Manager where record ownership cannot
be identified or is in dispute.
Assign the appropriate analyst for investigation and diagnosis
effort.
Refer Policy: Ownership, Notification, Escalation, Reject a
Record
Problem Manager Ensure record ownership is assigned to relevant support
group.
Perform notification when reassigning record ownership.
Manage any issues relating to record ownership.
Refer Policy: Ownership, Notification, Escalation
10 Support Group Assess whether investigation into the root cause of a problem
is justified and warranted at this point in time.
Perform Root Cause Analysis (RCA) now?
If Yes - go to step 13 Investigate and Diagnose
If No - go to step 11 Document Reason
Refer Policy: Root Cause
11 Support Group Document the reason(s) why Root Cause Analysis of a
problem is not justified and warranted at this point in time.
Notify the Problem Manager of the reason why the problem is
to remain unresolved.
Refer Policy: Notification, Unresolved Problems
Problem Manager Communicate to Incident Management any information
received relating to unresolved problems.
Gather and compile information on unresolved problems as
input to relevant Problem Management meeting.
Refer Policy: Meetings
Step Activity Role Activity Description
12 Problem Manager As a parallel activity using online searches and generated
reports, continuously monitor and track the progress of all
problems and known errors through to closure.
Identify and initiate any required actions to progress a problem
(i.e. re-categorising, reprioritising, correcting ownership or
escalating problems)
Ensure record is regularly updated with current status and the
actions performed.
Refer Policy: Monitoring Progress, Notification, Escalation
13 Continue within the Problem Management process by going to
the next activity: Investigate and Diagnose.

Table 3 Detect, Log, Categorise & Prioritise: Problem Activity Description


Investigate and Diagnose – Problem Activity Flow

Figure 3 Investigate and Diagnose: Problem Activity Flow


Investigate and Diagnose - Problem Activity Description
Step Activity Role Activity Description
1 Support Group Gather any additional data to commence problem investigation
Reference the knowledge base to identify any existing
workaround(s) for the problem.
Note: Data may be sourced from events, operational
monitoring, recent Change Management activities,
configuration item information
2 Support Group Based on the outcome of the investigation does a workaround
exist?
If Yes – update problem record with workaround information
and go to step 8 Determine Root Cause
If No – go to step 3 Create Workaround
Note: Limit the time and effort spent in finding or creating a
workaround as the primary focus is to identify the root cause
and provide a permanent solution.
3 Support Group Create an appropriate workaround where one does not exist.
Identify the appropriate steps or actions required.
Take into consideration any risks associated with the
workaround.
Engage other support groups to assist with workaround where
required.
4 Support Group Confirm that a workaround has been created.
Workaround created?
If Yes –go to step 7 Update Knowledge base
If No – go to step 5 Communicate workaround unavailable
5 Support Group Notify the Problem Manager to communicate where a
workaround for a problem does not exist.
Go to step 8 Determine Root Cause
Problem Manager Receive notification of an unavailable workaround.
Determine best course of action and engage other
stakeholders as required.
Gather and compile this information as input to any Problem
Management meeting.
6 Account Team Receive communication of any unavailable workaround for a
problem.
Where required, liaise with Problem Manager and ClientIT to
assist with determining the best course of action to be taken
ClientIT Receive communication of any unavailable workaround for a
problem.
Where required, liaise with Account Team to assist with
determining the best course of action to be taken.
7. Support Group Update the knowledge base with information pertaining to the
newly created workaround.
Update the problem record with the workaround information.
Communicate available workaround information to Incident
Management so that workaround can be applied to incidents
as required.
8 Support Group Perform root cause analysis to determine the underlying cause
of a problem.
Identify any contributing factors.
As required, engage others in parallel investigation to assist
with identifying underlying cause of a problem.
Note: Engage other Service Providers via Vendor.
Update problem record with the results and findings of the
analysis
Re-categorise problem where required.
Note: To assist with determining the root cause, a meeting
with key stakeholders may be required. This meeting is
facilitated by the Problem Manager.
Refer Policy: Ownership
Step Activity Role Activity Description
9 Support Group Based on the findings of the RCA, was the root cause of the
problem found?
If Yes – go to step 11 Create Known Error and
Communicate
If No –.go to step 10 Determine Course of Action
10 Support Group Document the results and findings, where root cause analysis
was completed and the underlying cause of the problem
remains unidentified.
Engage Problem Manager to assist with determining the best
course of action.
Execute the agreed action and update the problem record.
Problem Manager In consultation with Support Group determine what course of
action should be taken where the root cause is not identified.
Ensure problem record is updated with the determined course
of action.
Provide information of unresolved problem to Incident
Management and update knowledge base where required.
Gather and compile information on unresolved problems as
input to Problem Management meeting.
Refer Policy: Unresolved problem
Account Team Receive communication when root cause of a problem is
unidentified.
Where required, liaise with Problem Manager and ClientIT to
assist with determining the best course of action to be taken
ClientIT Receive communication when root cause of a problem is
unidentified.
Where required, liaise with Account Team to assist with
determining the best course of action to be taken.
11 Support Group Convert the problem into a known error where the root cause
has been identified and where possible a workaround exists.
Update the knowledge base and communicate to Incident
Management the known error information.
Gather and compile information on known errors as input to
Problem Management meeting
12 Problem Manager Provide Account Team with gathered problem information
relating to the current status of problems for review and
discussion at the ClientProblem Management meeting.
Refer chapter 6 Meetings
Account Team Facilitate the ClientProblem Management meeting with
ClientIT and key stakeholders

ClientIT Participate in the ClientProblem Management meeting.

13 Support Group Engage other analysts (including 3rd parties) to assist with the
investigation of a problem and the determination of a
workaround and root cause.
14 Support Group As a parallel activity using online searches and generated
reports, continuously monitor and track the progress of all
problems and known errors through to closure.
Identify and initiate any required actions to progress a problem
(i.e. re-categorising, reprioritising, correcting ownership or
escalating problems)
Ensure record is regularly updated with current status and the
actions performed.
Refer Policy: Monitoring Progress, Notification, Escalation
Continue within the Problem Management process by going to
the next activity: Resolve and Close Problem

Table 4 Investigate and Diagnose Problem Activity Description


Resolve and Close - Problem Activity Flow

Figure 4 Resolve and Close Problem Activity Flow


Resolve and Close - Problem Activity Description
Step Activity Role Activity Description
1 Support Group Based on the known error information, identify most appropriate
solution to resolve problem and eliminate the known error.
Where required, engage others including Service Providers and
3rd parties to contribute to the solution.
Reassign ownership where solution is to be provided by another
group.
Refer Policy: Ownership, Notification, Escalation
2 Support Group Create a plan for the identified solution.
Define the approach and outline all key resolution activities.
Where applicable include the removal of any previously applied
workaround.
Where a proposed solution is applicable across the IT
environment, log a proactive problem to initiate preventative
measures and thereby prevent future incidents.
3 Support Group Assess the proposed solution to determine its feasibility and
whether sufficient justification exists to resolve the known error.
Consult with key stakeholders in relation to risks, costs or
benefits from a business and technical perspective.
Consider exploring any other alternative solution options.
Verify the entitlement of the proposed solution and escalate to
Account Team any entitlement issues.
Refer Policy: Known Error
Problem Manager Consult with Support Group and other stakeholders in relation to
the feasibility and justification of the proposed solution.
Account Team Consult with ClientIT and other stakeholders in relation to the
feasibility and justification of the proposed solution.
Handle any entitlement issues and liaise with ClientIT to obtain
an agreed outcome and communicate to key stakeholders.
Ensure problem record is updated with agreed outcome.
ClientIT Consult with Account Team and other stakeholders in relation to
the feasibility and justification of the proposed solution.
Liaise with Account Team to resolve any entitlement issues.
4. Support Group Is the resolution of the problem justified?
If Yes – Formally document the solution in the problem record
and go to step 6 Initiate Request for Change (RFC)
If No – Go to step 5 Document Reason and Notify
Stakeholders
5 Support Group Document in the problem record the reason why the proposed
solution is not feasible or justifiable and a decision is not to
explore any other options.
Ensure record reflects where ClientIT agrees to accept any
ongoing risk.
Place the known error in the appropriate status to indicate it will
remain unresolved and notify all key stakeholders including
Incident Management.
Note: The documented workaround will be applied to any future
recurring incidents.
Refer Policy: Notification, Unresolved Problems
6 Support Group Where criteria met, initiate a Request for Change (RFC) with the
required information so as to execute the agreed resolution
plan.
Refer Process: Change Management
Step Activity Role Activity Description
7 Support Group Monitor the problem resolution via Change Management.
Apply the permanent solution as per the documented resolution
activities.
Ensure the appropriate technical and business testing is
performed to confirm whether the outcome of the change
implementation was successful in resolving the problem.

8 Support Group Is the problem confirmed as resolved?


If Yes – go to step 9 Complete Problem Record and Notify
If No – Update problem record to indicate the implemented
solution was unsuccessful and
 Reassess the solution by returning to Step 1 Search for
Resolution or
 Reinvestigate the root cause by returning to Activity:
Investigate and Diagnose
9 Support Group Complete the problem record by ensuring it contains the
following accurate, correct and detailed key information:
 Workaround
 Outcome of root cause analysis
 Resolution and verification details
 Any related record identifiers are referenced
 Categorisation of problem.
Update knowledge base with any problem and or known error
information.
Notify Problem Manager upon completion of problem record
Note: Do not reassign record ownership to Problem Manager
queue.
Refer Policy: Notification
10 Problem Manager Verify the quality of the record content.
Address any outstanding issues with the relevant Support
Group.
Close the problem record and perform notification to key
stakeholders where applicable.
Ensure all related incident and problem records are also closed.
Refer Policy: Notification, Close Record
11 Problem Manager Assess whether the closed problem is considered “major”.
Note: A Major problem is based on the duration of the problem’s
lifecycle being unusually prolonged due to a high number of
issues: not just where it’s allocated a high priority.
Is it a major problem?
If Yes – go to step 12 to Initiate Major Problem Review
If No – This process ends here. Continue with monitoring and
tracking progress of problems. Go to step 13 Monitor and
Track Progress
12 Problem Manager Where a closed problem is identified as major, initiate a review
as part of continuous improvement to determine:
 What was done right,
 What was done wrong
 What can be improved upon for next time
13 Problem Manager As a parallel activity using online searches and generated
reports, continuously monitor and track the progress of all
problems and known errors through to closure.
Identify and initiate any required actions to progress a problem
Ensure record is regularly updated with current status and the
actions performed.
Refer Policy: Monitoring Progress, Notification, Escalation
The process ends here.
Table 5 Resolve and Close Problem Activity Description
Appendix A: RCA Template
Root Cause Analysis Report
INCIDENT REPORT DETAILS - To be filled by service owners with assistance of Incident
Managers
IM#:
Incident Summary

Service Owner: NAME: ROLE:


Priority: Choose an item.
Date & Time of Occurrence: DATE: TIME:
Date & Time of Restoration: DATE: TIME:
Service Outage Duration:

Was this resolved within SLA? Choose an item.


If no, state the challenge.
Choose an item.
Did it require vendor support?
Detection and Monitoring
Mode of detection e.g. NOC, Alerts, Customer
complaint, logs, etc.

Communication
Did this require communication to the customers?
Was it done?

Was escalation to the Internal support teams or


vendors done in time?
Services impacted/Business Impact:
(Qualitative or Quantitative) List of impacted
services/affected nodes:
Sequence of Events

Date Time Key Event / Update Who

Root Cause Analysis with 5 whys:


Detailed Root Cause analysis of fault (attach supporting Logs and images and where necessary)
Why question #1
Answer#1
Why question #2
Answer#2
Why question #3
Answer#3
Why question #4
Answer#4
Why Question#5
Answer#5
Root Cause
Summary

Root Cause Classification APPLICATION/CONFIGURATION

If VENDOR/PARTNER, please specify.


Choose an item.
Choose an item.

Resolution details/corrective actions:


What was done to restore services
Incident recurrence:
Has a similar incident occurred in the past 3
months?
If yes, state number of occurrences and record the
IM numbers

NEXT STEPS (summary of Avoidance/Mitigation measures to be undertaken)


Action: Action Owner : Completion date: Status
i. Click or tap to enter a
date.
ii. Click or tap to enter a
date.
iii. Click or tap to enter a
date.
iv. Click or tap to enter a
date.

RCA Report Prepared by: ……………………………………………………………………….

Manager:…………………………………...…..

Department Head :………..……………...……………………………………

You might also like