Yale Problem Management Process Guide
Yale Problem Management Process Guide
Process Guide
Purpose
This document will serve as the official process of Problem Management for Yale University. This
document will introduce a Process Framework and will document the workflow, roles, procedures, and
policies needed to implement a high quality process and ensure that the processes are effective in
supporting the business. This document is a living document and should be analyzed and assessed on a
regular basis.
Scope
The scope of this document is to define the Problem Management Process, and process inputs from,
and outputs to, other process areas. Other service management areas are detailed in separate
documentation. This document includes the necessary components of the Process that have been
confirmed for the organization.
Role Description
Problem Management Process Ensures that all aspects of the problem management process are being executed
Owner effectively. The Problem Manager takes a quality assurance rule over problem
resolution teams and is responsible for assembling teams effectively.
Problem Owner Assigned a problem and uses the Problem Analysts, Subject Matter Experts and
others to help assess and resolve the assigned problem. In some cases, the
Problem Owner will also be the Service Owner. The problem record will be
assigned to the Problem Owner.
Problem Manager / Manages execution of the Problem Management process and coordinates all
Coordinator(s) activities required to respond to problems in compliance with SLAs and SLO's.
Receives problem candidates, assesses against criteria and initiates the problem
activities and eligible problems.
Service Owner Ensures the service is managed with a business focus, the definition of a single
point of accountability is absolutely essential to provide the level of attention
and focus required for its delivery.
The Service Owner is accountable for Continual improvement and the
authorization of changes and improvements to the service and has financial
accountability.
The following illustrates the Responsibility, Accountability, Consulted and Informed (RACI) matrix related
to the key Problem Management Activities:
Step Activities
1.1 Create Problem A problem can be triggered from a variety of sources. It generally is created from:
Record (candidate) Release Note or Vendor that states a known error
and submit for Major Incident where the root cause is unknown and problem investigation is requested by
consideration the Situation / Incident Manager
Proactive analysis of incident trend data
When a problem is detected, it is recorded with initial data and the rationale in considering the problem
for further investigation. This could include the association of incidents, log reports through event
management, written rationale from a Service Owner or notification from the developers or vendors
stating a known error is released into the environment.
1.2 Assess Problem Once a problem is detected it is assessed against pre-defined organizational criteria by the Problem
Record for Manager. During the initial stages of Problem Management, it may not be possible or logical to begin to
consideration work on ALL problems identified. There is some discretion the Problem Manager has to determine
which ones will be addressed in what order. This criteria may simply based on the Priority of the
Problem, ie only address Priority 1 Problems, or problems that have a minimum number of incidents
associated, etc. The Problem Manager may choose to obtain input from a Steering Committee that can
assist in committing resources.
By implementing Problem Management, effort required for Incident Management will decline and
resource effort will shift to Problem Management. Until PM is fully implemented and embraced in an
organization judgment may be required in identifying the problems to be addressed.
1.3 Meets Criteria? If the problem meets the criteria for continued investigation, move on to step 2.0 otherwise the
problem record can be closed with rationale and notification back to the initial detector.
Step Activities
2.1 Complete Problem Finalize all the coding and categorization (using same categorization as Incident) of the
categorization including results Problem making sure to associate all relevant incidents and include all relevant details,
from initial assessment including but not limited to:
User details
Problem details including service impacts, incident affects etc.
Equipment details
Priority and categorization
Associated Incidents
Details of all diagnostic or attempted recovery actions taken, for problem
team consideration to create as a formal KM record
Also include a description of how the problem met the criteria to become a recognized
problem.
Step Activities
3.1 Prioritize the Problem Update the Problem record with Impact and Complexity based on the Priority Matrix. Problem
prioritization is similar to Incident prioritization, however, it takes into account factors such as
costs, effort to resolve etc.
3.2 Identify and assign The Problem Manager determines who should be assigned ownership and assigns the Problem
Problem Owner to the appropriate Problem Owner (which is often the Service Owner).
3.3 Accept and take The Problem Owner receives the problem record and accepts ownership.
ownership of the Problem
3.4 Identify required skill The Problem Owner determines who else is required to participate as part of the resolution
sets for problem team. The Problem Manager, along with the Problem Analysts identified are the core problem
resolution team & obtain resolution team and SMEs can be called upon as needed to provide expertise during the
commitment investigation and resolution stage.
Identification is performed through the creation of tasks that can be assigned to assignment
groups for queue managers to identify a resource(s) to perform this activity. Assignees update
the tasks as their investigation activities continue. The problem owner updates the problem
task when/if required.
3.5 Resources Secured? The Problem Owner seeks the required individual involvement and if successful the problem
can continue on to the Investigation stage. If the Problem Owner is not successful in obtain
resources the Problem Owner may need to work with the Problem Manager to escalate within
the organization.
3.6 Work to obtain The Problem Owner and Problem Manager may need to consider alternatives if desired
necessary resource resources are unavailable. For example, they may seek alternate SMEs, delay investigative
involvement activities, escalate to senior management, etc.
4.1 Perform Root Cause Analysis The resolution team, made up of Problem Analysts (and calling on SMEs as
required), work at determining the root cause of the problem. The appropriate
level of resource and attention is determined by the priority of the problem.
Along with various problem solving techniques and tapping into available
knowledge such as Known Error Database can help to pinpoint the point of failure.
This step kicks off 2 potential questions… Is the root cause known and is there a
workaround available. A workaround may not be known but root cause is known
whereas a workaround may exist without knowing the root cause. These 2
streams can be done simultaneously.
4.2 Root cause known? Once the root cause is known, proceed to the “known error creation” step.
4.3 Workaround? If during the investigation and diagnosis stage a workaround is identified, process
to the “workaround creation” step.
4.4 Proceed with investigation? At any point in the process it may be determined that further investigation is not
required. This decision is made by the Problem Owner in consultation with a
variety of stakeholders including the Service Owner, SMEs, etc. It may be
determined that the effort involved in further investigation does not out-weigh the
benefit from resolving the problem. If the decision is made to NOT proceed with
problem activities, it must be clearly noted in the problem record.
Step Activities
5.1 Document Workaround Upon identifying a workaround, the problem resolution team documents the workaround
and submit for authorization for use by Service Desk and potentially by end users.
5.2 Authorize Workaround The Problem Owner, who is ultimately responsible for the resolution of the problem,
authorizes the workaround.
5.3 Deploy Workaround Upon authorization, the workaround is deployed to the appropriate levels in the
organization.
5.4 Root cause known? Although a workaround is identified, the Root cause may still not be known. If root cause is
known proceed to the “Known Error Creation” step and if it isn’t known, go back to the
“investigation and diagnosis” step.
Step Activities
6.1 Document Known Error and Upon identifying the root cause, the problem resolution team documents the
submit for approval known error in assigned tasks. The problem owner reflects the consolidated root
cause details in the problem record.
It is important to note that occasionally Known errors are identified with new
applications or by vendors, it is important to ensure these are recorded. This
allows a mechanism to track how often these known errors are being
encountered and will help to form the case towards a resolution.
6.2 Approve Known Error The Problem Owner, who is ultimately responsible for the resolution of the
problem, approves the known error.
6.3 Update Knowledge Record and Upon approval, the known error workaround is deployed to the appropriate
make available to Service Desk levels in the organization.
Step Activities
7.1 Work to resolve problem The problem resolution team works to identify temporary or permanent solutions
or potentially alternatives to be considered by the Service Owner. The solution
alternatives and options should be documented so the Service Owner has all the
information necessary to make an information decision on the course of action.
7.2 Validate solution and/or options The Problem Owner works with the resolution team to validate the solution
options being put forward for approval.
7.3 Assess solution alternatives and The Problem Owner and Service Owner discuss the available options to
approve course of action determine the best course of action. In some cases the Service Owner may need
to bring the decision to another decision-making body usually if additional
funding is required.
7.4 Solution Approved If the solution is approved, it will proceed through the Change Process. If the
solution is not approved a decision is made to continue investigation to come up
with an alternative solution or to discontinue efforts.
7.5 Problem Resolved? Once the solution is implemented, the Problem Owner together with the Problem
Analysts determine if the solution did in fact resolve the problem. If it did,
proceed to close the problem. If it did not, a decision is made to continue
investigation to come up with an alternative solution or to discontinue efforts.
7.6 Proceed with Investigation? This decision is made by the Problem Owner in consultation with a variety of
stakeholders including the Service Owner, SMEs, etc. It may be determined that
the effort involved in further investigation does not out-weigh the benefit from
resolving the problem. If the decision is made to NOT proceed with problem
activities, it must be clearly noted in the problem record.
Step Activities
8.1 Update Problem Once the Problem has been resolved or it has been decided that problem activities are to not
Record with sufficient continue, the problem record is closed.
coding & detail Note: This requires an opportunity to confirm that the problem has truly been addressed
through the change, and may require an extended timeframe to validate (e.g. monthly batch
processing may have to occur to be certain the change addressed the problem).
Part of closing the problem record is to ensure all approvals are documented, the coding is
accurate, any rationale for decisions are document.
8.2 Major Problem? Once the problem has been officially closed, if it is a Major Problem, it requires a Major Problem
Review.
Release and Deployment Management •Acceptable known errors captured during release review.
General Root Cause Request •Typical problem manager trend analysis activities.
Problem Types
Level 1 Level 2 Level 3 Descriptions
Reactive Trend Consistent The problem is based on a clearly identified and recurring
associated incidents trend.
Reactive Trend Inconsistent The problem is based on a trend of associated incidents that is
inconsistent but recurring.
Reactive One-Time Authorized The problem is related to incidents generated from a suspected
Change authorized change.
Reactive One-Time Un-Authorized The problem is related to incidents generated from a suspected
Change unauthorized change.
Reactive One-Time Major Incident The problem is related to a major incident where root cause
analysis was requested directly or determined to be necessary in
the major incident review.
Reactive Other The problem is related to one or more incidents that share some
other characteristic(s).
Proactive Release Pre- The accepted known error was identified as part of release and
Deployment deployment review activities. This also includes known errors
Known Error identified for COTS packages (i.e. release notes).
Proactive Event-Driven The problem is related to event monitoring warnings where the
(Warning) service has not yet been impacted from a customer’s
perspective.
Proactive Other The problem has been identified proactively through some other
means.
Problems often incur costs, either directly or through the assignment of critical support resources to
perform diagnosis and resolution activities. In addition, there is significantly more subjectivity in the
prioritization of problems vs. incidents, due to the nature of the process itself. The goal is to remove
impacts to customers and thus, there are often competing factors that must be considered including
cost to the business in lost revenue, technical complexity of the problem, relative impacts to customers
and costs to determine root cause (direct costs such as licenses/hardware etc., or as a result of
committing highly skilled, and therefore high cost resources to the problem teams).
Prioritization Matrix
High 3 2 1 N/A
Impact
Medium 4 3 2 N/A
Low 5 4 3 N/A
Complexity
Impact Values
Value Description
High The problem is causing a high number of customer impacts, often derived through the
volume and priority (e.g. high impact) of associated incidents. In addition, problems that are
deemed to be incurring high expense or lost revenue would be considered high impact.
Medium The problem is causing a some customer impacts, often derived through the volume and
priority (e.g. medium impact) of associated incidents. In addition, problems that are deemed
to be incurring expenses or potentially lost revenue would be considered medium impact.
Low The problem is having a minimal impact on customers, often derived through the volume and
priority (e.g. low impact) of associated incidents. No appreciable revenue lost is predicted.
Value Description
High The problem is complex due to factors including very high costs and/or significant effort
required by IT support staff to diagnose and/or remove the problem.
Medium The problem presents some complexity due to a combination of cost and/or requirement to
focus a large number of resources (or a selct few who are critical) to diagnose and/or
remove the problem.
Low Acceptable or minimum complexity due to costs and/or resource requirements to diagnose
and/or remove the problem.
Accepted Known Error - Workaround •The problem will not be removed as the workaround is
Implemented acceptable.
Accepted Known Error - No •The problem will not be removed and no workaround exists
Workaround however the impacts are minimal/acceptable.
•Used when a feature request has been raised, but the cost of
Unresolved – Cost the request is too high to action and acceptable to the
business/customer (payer).
•The feature request was already defined for a future release.
Unresolved - Future Release Unresolved problem may be associated to an originating
problem for the initial request.
Rejected
Logged by
In Known Pendin Root
New Draft Problem Progres Error g Cause Closed
s
m
Chang
e
Unresolved /
Resolved – No Action Taken /
Accepted Known
Deferred
RCA
Cause
Ref #:4 % Repeat Incidents Consider replacing with Incident to Problem Ratio?
Ref #:9 % of problems not resolved within Recommend against SLA reporting for PM
SLA targets
Ref #:11 % problems reopened Problems are not reopened – they are only resolved with problem removal is
confirmed
Ref #:12 % of problems with customer Problems associated to incidents / Incident to Problem Ratio? (Exclude
Impact incidents resolved with workarounds)
Ref #:14 % Problems responded on time Recommend against SLA reporting for PM
Ref #:15 % Problems resolved on time Recommend against SLA reporting for PM
Ref #:16 Cost of solving a problem Cost to resolve captured – calculate in QuickBase?
Ref #:7 Average problem resolution time Mean time to close problem
Ref #:10 % of problems not linked to Known Volume of problems with Known Error Flag
Errors records
Ref #:13 % Ageing problems Volume of undiagnosed (i.e. no known error) problems – backlog
Ref #:17 Total number of problems caused Volume of problems by Incident Source = unauthorized change
due to unauthorized changes
Ref #:18 Problems not associated with Volume of problems with no associated incidents
incidents