Problem Management Process Overview
Problem Management Process Overview
Table of Contents
Document Approver
Document Reviewer(s)
The following will trigger an exit from the Problem Management process:
Appropriate actions taken to resolve a problem to prevent future incidents and
problems and,
The activity of closing the problem record and a Major Problem Review initiated
where criteria met.
The scope of the Problem Management process includes a series of activities which
involve the following:
Categorisation and prioritisation of problems
Investigation (root cause analysis) into underlying cause and identification
Progressing problems into known errors where underlying root cause identified and
workaround available
Solution (and workaround) definition and selection
Where justified, resolving known errors by applying a permanent solution to prevent
any recurrence
Appropriate prioritisation of resources required for resolution based on business
need
Submission of change requests to implement solutions
Continuous monitoring and tracking of problems and known errors (taking
appropriate action where required)
Communicating problem information.
Contribution to the collective problem resolution knowledge base and solution
creation
Formulating recommendations for improvement, maintenance of problem records
and review of the corrective actions
Updating records with problem and known error information (KEDB).
It is vital that all information entered in to the IT Service Management tool is accurate and follows the
guidelines set in this document.
Upon the detection of a problem either proactively or reactively, Problem Management process must
be invoked immediately to log the problem with the following minimum information:
Unique Problem record identifier
Date and Time Problem recorded
Description of problem and any actions already taken to assist with investigation
Priority as per policy
Configuration item(s) involved
Classification (category)
Indication, where problem has been proactively identified or is a recurrence.
Any Third Party record number for referencing purposes
Where a problem is reactively detected, a relationship must be built by linking the logged
problem to all associated incidents.
To log a problem in the IT Service Management tool, one must have the appropriate tool access.
Procedure to log a Problem Ticket is explained in the attachment as Appendix B on Page 30.
Priority
The priority of the problem is inherited from the associated incident. Ref priority Policy in
incident management process for the priority of incident
Ownership
Ownership is established by assigning a record to the appropriate support group regardless of formal
acceptance or acknowledgment.
There is only one owner of a record at any given time, and that is the support group it is currently
assigned to.
Any reassignment of a record may occur throughout the Problem Management process.
Prior to any record reassignment the owning support group ensures that the reason and all actions
leading up to the reassignment have been documented in the record.
Perform verbal notification upon record assignment to ensure awareness and no delays are
experienced.
Where ownership disputes cannot be resolved, immediate escalation to the Problem Manager must
occur.
Parallel Activities
Other Support Groups may be engaged to assist and perform parallel activities where record
ownership resides with another support group.
Irrespective of record ownership, where parallel activities are required, a problem is to be treated
based on its allocated priority.
Where a Support Group is engaged in parallel activities, it is their responsibility to document all
investigation activities and findings in the problem record.
Where a Third Party is engaged in parallel activities and does not have direct IT Service Management
tool access, then the responsibility for ensuring updates to the problem record remains with a support
group that owns the relationship (i.e. acting on behalf of the Third Party).
Where a request for parallel activities is refused, immediate escalation to the Problem Manager must
occur.
Reject a Problem
When ownership is rejected by the assigned Support Group, the record must be
updated to indicate the reject reason and then either:
Reassign the problem to the correct Support Group (where known) or
Contact Problem Manager to establish an appropriate owner
Verbal notification must accompany any reassignment to ensure no further delays are
experienced.
Where ownership disputes occur, escalate to the Problem Manager.
Notification
Notification occurs throughout this Problem Management process. This ensures that the
responsible role is notified about further action required by them, or to communicate
relevant information on a problem’s progress.
All notifications performed including any delays experienced must be recorded in the
problem record.
The table below outlines the key notifications by the various roles involved in the Problem
Management process.
The ClientIT and any other Service Providers upon receiving notification from Vendor are responsible
for communicating internally within own organization.
If the The Notifies the By
Analysis highlights a service Support Group Problem Manager E-mail communication to
improvement opportunity central mailbox Vendor’s
Problem Mngt Team
Priority of a Problem requires Support Group Problem Manager Verbal notification to obtain
alteration prior approval
If the The Notifies the By
Record ownership is Rejected Existing Support Support Group to Verbal notification upon
and correct ownership is Group which the record is rejecting ownership
known being reassigned Record updated with reject
reason
Existing Support Problem Manager Verbal notification upon
Group rejecting ownership
Record updated with reject
reason
Workaround exists for a Support Group Problem Manager Record updated
problem E-mail communication
Problem Manager Stakeholders Verbal notification and
email notification
Known Error solution requires Support Group Other Support Group Verbal and email
involvement from other and managed 3rd communication
analysts or third parties Party Suppliers Record updated
Proposed solution requires Support Group Account Team and E-mail notification and
consultation with ClientIT Problem Manager Record updated
Account Team ClientIT Verbal notification
Escalation
Escalation occurs throughout the Problem Management process. This ensures additional
attention and visibility is given to issues so as to meet service level agreements and or
customer expectations.
Escalation may be functional or hierarchic and the level of escalation is dependent on the
nature and significance of the issue.
All escalations must continue until an agreed outcome is achieved. The outcome of the
escalation must be documented in the problem record and communicated accordingly.
The table below outlines the key escalations by the various process roles involved in the
Problem Management process.
Root Cause
A support group must conduct a Root Cause Analysis (RCA) for each problem
regardless of priority.
Where the cause of a problem cannot be identified, the assigned Support Group is to
establish preventative measures so as to capture additional data to assist with root
cause identification or prevent future incidents.
RCA need to have its scope ( In or out of Lycamoney) defined specifying severity (S1or
S2) as well
The Problem Manager will review and accept the results prior to communicating the
information to relevant stakeholders.
RCA for the Priority 2 problem record must be filled in the tool only, having Symptom,
Cause, Resolution details and all the identified corrective and preventive actions
(CAPA). The identified actions items must be tracked in the ITSM tool.
An exemption from both the Problem Manager and Account team must be obtained
where RCA will not be performed, as it’s deemed as not warranted or justified. All
relevant information (i.e. the reason, exemption granted or not) must be documented in
the problem record.
Unresolved Problems
A problem remains unresolved where:
The root cause of a problem was investigated but the underlying cause was not
identified.
Analysis into the root cause of the problem was not warranted or justified based on
the impact and urgency assessment and incident count(s).
Root cause is identified and solution proposed, however an agreement has been
reached between the Account Team and ClientIT not to resolve the known error.
Where a problem remains unresolved a record must not be closed so as to ensure any future
incidents are identified, matched and tracked against the associated problem.
Unresolved problem records that have remained open for a defined period of time (i.e. 13 months) will
be reassessed by the Problem Manager
The Problem Manager must document the assessment outcome in the problem record
and close the record only after agreement reached post discussion with Account Team.
Known Error
A problem progresses to a known error once a root cause and workaround has been successfully
diagnosed and documented.
To resolve and eliminate a known error, an appropriate solution must be identified. This
solution is evaluated as to whether eliminating the known error is feasible and justifiable.
Prior to any implementation, a solution must be reviewed and approved by both technical
and operation stakeholders. The approved solution activities are formally documented
and implemented in accordance with the Change Management process (where criteria
met).
Where the solution is not to resolve and eliminate a known error, document the reason
and the agreement obtained within the account and operations team
Close a Record
Problem Manager must review and approve the resolution of all problem records. All approvals must
be documented in the record prior to closure.
Once a record has been resolved any recurrence of the same or similar problem must be recorded as
a brand new problem and a relationship established.
A problem record may be cancelled only where it is was opened in error and no problem activities
have commenced. Prior to cancellation provide cancellation reason in the record.
Monitoring Progress
Problem Manager must actively monitor the progress of all problems and known errors to
ensure their progression within lifecycle of the Problem Management process.
Specific attention must focus on those problems and known errors having an extended
period of inactivity or slow progression. Liaise with relevant party (i.e. Support Group,
Account Team) and initiate any required actions so as to progress the problem or known
error.
Problem Manager must ensure that the problem record is regularly and accurately
updated throughout its lifecycle.
Interfacing Processes
Problem Management process interfaces with the following processes:
1. Incident Management provides incident data for analysis so as to identify trends,
and to group related incidents that will assist in identifying whether a problem
already exists or requires creation. Problem Management provides information
relating to unresolved problems, Known Errors, available workarounds and
problem closure.
2. Problem Management will initiate a Request for Change (RFC) where it is
identified that a configuration item is to be altered in relation to a workaround,
permanent solution or preventative measure. Change Management provides the
status and the final outcome (i.e. successful or unsuccessful) of the implemented
change.
3. Run and Monitor Operation (Event Management) provides Problem
Management with real time and historical event information to assist with the
proactive detection of problem(s), the investigation of a problem and service
quality improvement.
4. Configuration Management provides Problem Management with configuration
item information to assist with problem investigation, resolution and preventative
measures. Problem Management provides information and modifications to
existing configuration items to enable the configuration data to be validated and
updated.
5. Key availability data is provided to Problem Management so that investigation,
resolution and preventative actions can be initiated to enable availability targets
to be met. Problem Management provides SLA Management Process with
problem information to allow for availability trends to be identified and remedial
action to be instigated.
Capacity information is provided to assist with problem investigation, resolution and pro-
active problem management activities. Problem information is provided to Capacity
Management to assist with capacity planning
3. Roles and Responsibilities
Responsibilities have been defined in relation to the Problem Management process activities using a
RACI model. In the table below are the defined RACI authority values.
R Responsible Individual(s) that are responsible for performing and completing a specific
activity.
The degree of responsibility is defined by the Accountable person.
(Note: Responsibilities may be shared)
A Accountable An individual who has the prime lead and the ownership and is ultimately
accountable for ensuring a specific activity is completed.
C Consulted Individual(s) who are consulted prior to a final decision or action.
(Two-way communication)
I Informed Individual(s) that receive information after a decision or action is taken. (One-
way communication)
Table 1 RACI Model
In the Roles and Responsibility matrix below, each responsibility has been defined
against functional role(s) with a defined RACI authority value.
Problem Manager
No. Responsibility Description
Support Group
Account Team
(Vendor)
(Vendor)
(Vendor)
ClientIT
Responsibilities for Detect, Log, Categorise and Prioritise
1 Re-actively detects a problem based on analysis of incident data R R,A
2 Proactively detects a problem based on analysis of information from R A
various other sources
3 Identify incidents not matched to problems or known errors R A
4 Identify and provide input into any service improvement opportunities R R R,A
5 Where a problem does not already exist, log a problem in the IT Service R R,A
Management tool with required information
6 Build relationships (link) the problem to associated incident(s) R R,A
7 Classify problem to determine category R R,A
8 Allocate the priority based on the impact and urgency assessment taking R R,A
into consideration the severity (seriousness) of problem
9 Identify and assign problem to the appropriate support group R R,A
10 Escalate where ownership cannot be determined or disputes exist R R,C A,C,
I
11 Assess whether investigation into the root cause of a problem is R R A
warranted or justified
12 Document the reason and communicate where root cause analysis is not R R,A C,I C,I
performed and the problem remains unresolved
Responsibilities for Investigate and Diagnose
13 Confirm ownership or reroute record where required R A
14 Gather additional information to commence investigation R,A
15 Obtain prior approval when altering the Priority of problem R A C C
16 Search for an available workaround, or create a workaround where one R,A
does not exist
17 Communicate workaround information R A,I I I
18 Update knowledge base with problem information R,A
19 Perform Root Cause Analysis to determine cause, engaging others in R R,A
parallel investigation
20 Notify where service levels are at risk R R,A I I
21 Formally document the findings of the RCA investigation R,A
22 Determine course of action where root cause is unidentified R A,C C
23 Convert problem into known error where cause is identified R A
Problem Manager
No. Responsibility Description
Support Group
Account Team
(Vendor)
(Vendor)
(Vendor)
ClientIT
24 Provide Incident Management (and Service Desk) with problem and R R,A
known error information
Responsibilities for Resolve and Close Problem
25 Diagnose known error and identify appropriate solution R,A
26 Develop a resolution plan outlining the key activities based on the R,A
proposed solution
27 Verify proposed solution with key stakeholders to determine the solution’s R A C C
feasibility and whether sufficient justification exists to resolve problem
28 Raise a proactive problem where the solution applies to the rest of the R,A
environment so as to prevent future incidents
29 Document reason where proposed solution is not feasible or justifiable R,A A C
and problem remains unresolved, handling any entitlement failures
30 Initiate Request for Change (RFC) or submit project proposals to R,A
implement agreed solution
31 Apply permanent solution as per the documented resolution activities R,A
32 Confirm elimination of Known Error by testing the applied resolution R,A
33 Complete problem record by validating that all key information was R,A
documented, adjust problem category where required
34 Update knowledge base with problem or known error information R,A
35 Close problem record as per policy and any other related problem records R,A I I
and perform communication of closure
36 Notify Incident Management to ensure all related incidents are also closed R,A
37 Initiate an internal Major Problem Review as per policy R,A
38 Obtain any required approval prior to closure of an unresolved problem R,A C C
39 Perform notification where Known Error remains unresolved R,A I I
Responsibilities across Problem Management activities
40 Establish and produce reports to communicate problem information R R,A
41 Provide regular progress updates in SMT record through to closure R A
42 Attend problem review meetings where required R R,A R
43 Facilitate customer meeting and distribute customised report as an out R,A I
from any problem review meetings
44 Manage record ownership disputes and issues R A,C
45 Continuously monitor and track the progress of problems and Known R,A
Errors through to closure
46 Identify problem records not progressing and initiate required actions , C R,A
communicating where required
Table 2 RACI Matrix
4. Key Performance Indicators (as per contract or mutually agreed
target)
[Link]. KPI description Measurement
RCA document Submission within 5 business days
75% *
1 timelines – for Severity 1 problem ticket
4 Support Group Log the detected problem in the IT Service Management tool
with the required information.
Refer Policy: Log a Problem
Problem Manager Log the detected problem in the IT Service Management tool
with the required information.
Refer Policy: Log a Problem
5 Support Group Build a relationship by linking matched incidents to an existing
problem or newly logged problem record.
Problem Manager Build a relationship by linking matched incidents to an existing
problem or newly logged problem record
6 Support Group Based on the available information, classify the problem into
its correct category (i.e. Hardware, Software, Application)
Note: A problem may be re-categorised at any stage
throughout the Problem Management process.
Problem Manager Based on the available information, classify the problem into
Step Activity Role Activity Description
its correct category (i.e. Hardware, Software, Application)
Note: A problem may be re-categorised at any stage
throughout the Problem Management process.
7 Support Group Determine the impact of the problem by performing an
assessment as to what effect a problem has or may have on
the business and any service levels.
Determine the urgency of the problem by assessing how long
the business can tolerate the unresolved problem.
Take into consideration the severity and seriousness of the
problem from an IT infrastructure perspective.
As required, liaise with Problem Manager and or Account
Team to determine the correct impact and urgency of a
problem.
Refer Policy: Priority
Problem Manager Determine the impact of the problem by performing an
assessment as to what effect a problem has or may have on
the business and any service levels.
Determine the urgency of the problem by assessing how long
the business can tolerate the unresolved problem.
Take into consideration the severity and seriousness of the
problem from an IT infrastructure perspective.
As required, liaise with Account Team and Support Group to
determine the correct impact and urgency of a problem.
Refer Policy: Priority
Account Team As required, liaise with Support Group, Problem Manager and
ClientIT to determine the correct impact and urgency of a
problem.
ClientIT As required, confirm business impact and urgency in
consultation with Account team.
13 Support Group Engage other analysts (including 3rd parties) to assist with the
investigation of a problem and the determination of a
workaround and root cause.
14 Support Group As a parallel activity using online searches and generated
reports, continuously monitor and track the progress of all
problems and known errors through to closure.
Identify and initiate any required actions to progress a problem
(i.e. re-categorising, reprioritising, correcting ownership or
escalating problems)
Ensure record is regularly updated with current status and the
actions performed.
Refer Policy: Monitoring Progress, Notification, Escalation
Continue within the Problem Management process by going to
the next activity: Resolve and Close Problem
Communication
Did this require communication to the customers?
Was it done?
Manager:…………………………………...…..