0% found this document useful (0 votes)
455 views79 pages

ITIL 4 Problem Management Guide

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
455 views79 pages

ITIL 4 Problem Management Guide

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Professionals Organisations Partners

dL
< Back

ITIL® 4 Problem Management Official


Practice Guide
April 14, 2023

39 min read Practitioner Resources ITIL4 Practice Guides ITIL

  169 Likes
This guide provides practical guidance for the problem management
practice.

1. About this guide

2. General information

3. Value streams and processes

4. Organizations and people


5. Information and technology

6. Partners and suppliers

7. Capability assessment and development

8. Recommendations for practice success

9. Glossary

10. Acknowledgements

1. About this guide


This guide provides practical guidance for the problem management practice.
It is split into seven main sections, covering:
· general information about the practice
· the practice’s processes and activities and their roles in the service
value chain
· the organizations and people involved in the practice
· the information and technology supporting the practice

· considerations for partners and suppliers for the practice


· information on assessing and developing the capability of the
practice
· recommendations for succeeding in the practice.

ITIL® 4 qualification scheme


Selected content from this guide is examinable as a part of the following
syllabi:
· ITIL® 4 Specialist: Create, Deliver, and Support

· ITIL® 4 Practitioner: Problem Management


· ITIL® 4 Specialist: Monitor, Support, and Fulfil
Please refer to the respective syllabus documents for details.

2. General information
2.1 Purpose and description

Key message
The purpose of the problem management practice is to reduce
the likelihood and impact of incidents by identifying actual and
potential causes of incidents, and managing workarounds and
known errors.

No product or service is perfect. Every product has errors which can cause
incidents. Errors may originate in any of the four dimensions of service
management. For example, a mistake in a third- party contract is as likely to
cause an incident as a technical component failure. Many errors are identified
and resolved during design, development, or testing before a product goes
live, and don’t affect the service. However, some errors will remain
undiscovered and proceed to the live environment, and these may cause
incidents.
The increasing complexity of the business and technology environment adds
to uncertainty. As more combinations of digital and business products,
services, and workflows are created, there is a greater chance of these leading
to incidents, despite each individual element being well-known, thoroughly
tested, and safe-to-fail.
The problem management practice is adopted and developed by service
providers to ensure that errors in the live environm ent are identified, analysed,
and, where required and possible, removed or fixed.
This practice is beneficial for both IT service providers and their service
consumers. Benefits for service providers include:
· Increased reliability of IT services

· Reduced losses and costs caused by IT service unavailability or


degradation
· Fulfilment of the service quality targets
· Reduced technical debt

· More even and predictable utilization of IT support resources.


Benefits for service consumers include:
· Increased reliability of business operations and business services

· Reduced business risks


· Reduced losses caused by business service unavailability

· Better image due to uninterrupted business services.

2.2 Teams and concepts


Errors that may cause (or have already caused) incidents are called problems.

Definition: Problem

A cause, or potential cause, of one or more incidents.

Problem management has three distinct phases, shown in Figure 2.1.


Figure 2.1 The three phases of the problem management practice

The key features of each phase are described in Table 2.1


Table 2.1 Key features of the problem management phases

What is Purpose of the Output of the Performed by


known at the phase phase
beginning of
the phase

Problem There are Assessment of Registered, Many teams,


identification incidents potential initially including
which are impact of the described and technical,
likely to have errors and classified support,
a common, vulnerabilities, problems operations,
currently to decide development,
unknown whether they supplier
underlying are worth an managers and
error investigation others
or
There are
vulnerabilities
in our
products and
services,
which may
cause
incidents
Problem Registered, Investigation Discovered Technical
control initially and analysis of and assessed teams
described and the problems, errors in or
classified to understand products and Product/service
problems their current services, teams
and potential assigned to
impact. their owners Often multiple
teams working
Identification of together
the most
suitable owner
for each error

Error control Discovered Identification of Problem Assigned


and assessed the best course solutions technical
errors in of action for the (systemic teams or
products and known errors. and/or specialists
services, workarounds)
assigned to Planning and or
their owners implementation Solutions for
of problem incidents
solutions caused by the
or problems
periodic control
of the known Solved
errors; planning, problems
communication
and
implementation
of the problem
workarounds.

Not every problem goes through all three phases. Some can be dismissed at
the identification phase (as duplicates, or due to low impact on the products
and services). Others may be closed at the problem control stage because
their impact has changed during the investigation, and they need no further
attention.
Each phase has its own timeline.
Problem identification is a relatively quick workflow, although it can be
preceded by a long period of data collection and processing.
In the problem control phase, service providers can influence the speed of
investigation by assigning additional resources or by prioritizing the
investigation over other tasks. However, the speed is also influenced by the
difficulty of the investigation. Problem control can take minutes, hours, days, or
weeks, though service providers usually try to limit the time of investigation to
keep it cost efficient.
The timeline of error control timeline can vary significantly. When there is a
known problem solution to be implemented, the implementation can be
planned with high certainty, but still may take a long time. If there is no
reasonable way to resolve the problem, it may remain open for long time,
especially if there are effective ways to mitigate its impact. In this case the
problem remains open and undergoes periodic reviews to confirm or change
the course of action.
2.2.1 Problem identification

There are two main approaches to problem identification: reactive or


proactive.
A reactive approach is focussed on investigating the causes of incidents that
have already happened. This approach begins by analysing the symptoms and
then proceeds to the causes. It aims to prevent incidents recurring, and may
also contribute to the resolution of open incidents.
A proactive approach is focussed on identifying problems before they cause
incidents, assessing the related risks, and optimizing the response with the
aim of minimizing the probability and/or the impact of incidents. Proactive
problem identification and is based on information about vulnerabilities and
errors in the live environment which became known from sources other than
incident management. The information sources include:
· vendors providing information on vulnerabilities in their products
· developers, designers, or testers discovering errors in live versions
while working with subsequent versions
· user and specialist communities sharing their experiences of other
organizations
· monitoring of the infrastructure discovering deviations in systems
performance that do not yet qualify as incidents
· technical audits and other assessments.
Key message
Reactive or proactive?
Problem identification is always reactive to problems: it
does not prevent them from occurring the first time. The
proactive/reactive distinction refers to how problem
identification and investigation relates to incidents:
· proactive problem management helps to
prevent incidents from occurring the first time.
· reactive problem management helps to
prevent incidents from recurring and may help
to resolve open incidents.

Problem identification leads to the registration of a problem record. The


problem records include information on the initial categorization and initial
assessment of the business impact and urgency.
The purpose of problem categorization is to identify the team or specialist who
will be responsible for the problem investigation. Therefore, categories usually
reflect how the IT service provider teams are organized. In product-focused
organizations, where teams with mixed technical expertise are formed around
products, categorization may be based on the product or service catalogue. In
more traditional organizations with teams formed based on the technical
domain (applications, networks, databases, and so on), categorization is also
likely to be based on the technical domains.
Some problems occur outside of technical domains and may be related to
other dimensions of service management:
· value streams and processes (errors in procedures, processes, value
streams)
· organizations and people (lack of knowledge, unclear
responsibilities, errors in incentives schemes)
· partners and suppliers (errors in contracts, unclear responsibilities,
lack of due diligence).
It is important to remember that initial categorization is likely to change at the
problem control phase: many problems are found in overlaps of different
technical domains, or different products and services. The purpose of
categorization is to identify the team most suitable for leading the
investigation, not performing 100% of the related activities.
Impact and urgency are important inputs for problem prioritization. The initial
assessment of the business impact and urgency will differ for problems that
are identified proactively or reactively.
The business impact of problems that are identified through available
information about incidents can be estimated based on:
· the individual impact of the incidents
· the number and frequency of the incidents
· trends in the occurrence of incidents

· the expected change of the impact due to business cycles (for


example, seasonal business activities).
The urgency of the problem investigation in such cases depends on the
availability and effectiveness of incident resolutions. If the incidents have no
effective resolution, and the problem needs to be investigated for the
incidents to be solved, the urgency of the problem investigation is based on
the agreed time to resolve the incidents, and the problem investigation will be
initially focused on finding the cause(s) of the incidents and an acceptable
resolution for those incidents. If the incidents have an effective resolution, and
their impact is effectively minimized through the incident management
practice, the urgency of the problem investigation is lower, and this may lower
the problem priority at the problem control phase.
The business impact of problems identified proactively is estimated based on
the expected effect of the problem on IT and business services. This requires a
good understanding of the business and IT architecture and dependencies of
the IT products and services on the components identified as bearing an error.
The urgency of these problems is estimated based on the forecasted
probability and proximity of incident occurrence, and the possible business
impact of these incidents.
A detailed description of the problem identification processes is provided in
sections 3.1.1 and 3.1.2.
2.2.2 Problem control
Registered problems are accepted for analysis based on their initial
categorization, business impact and urgency.
If a problem has been assigned to a team which has sufficient resources to
start working on the investigation without delays, the team should begin the
problem investigation. Details of the problem control process are provided in
section 3.1.3.
However, the team may have other tasks waiting in a backlog, including other
problems identified and assigned earlier. Resources, on the other hand, are
always limited, and so is the number of tasks the teams can perform
simultaneously. This is where prioritization is needed.

Definitions

Prioritization
The action of selecting which tasks to work on first when it is impossible to assign resources to all
tasks in the backlog.

Task priority
The importance of a task relative to other tasks. Tasks with a higher priority should be worked on
first. Priority is defined in the context of all tasks in a backlog.

Prioritization is a tool for assigning tasks to people in the context of a team. If


an incident is handled by multiple teams, it will be prioritized within each
team depending on resource availability, target resolution time, and estimated
processing time. If several tasks need to be performed by different teams
working in parallel, each team will be prioritizing their own task:
· Prioritization is needed only when there is a resource conflict.
Where there are sufficient resources to process every task within
the time constraints, prioritization is unnecessary.
· In each team, all types of tasks (including problems) should await
prioritization and assignment in a single backlog, together with
other tasks (planned and unplanned).
· Visualization tools, such as Kanban, and Lean principles, such as the
limiting of work in progress, are useful for effective prioritization.
The following are some specific guidelines on the prioritization of problems:
· The target completion time for problem control is defined based on
the problem impact and urgency assessed at the problem
identification phase. They may be reassessed as new relevant
information is discovered by problem investigation.
· The target completion time for problem control is different from
the target resolution time. The target resolution time for problems
can only be planned after the problem control phase is completed,
as it is based on the understanding of the error(s) and available
ways to fix them. It is also possible that the problem control phase
concludes that there is no need or reasonable way to fix the
problem.
· Problems should be prioritized within a single backlog which
includes all tasks assigned to the team.
· Problem investigation may require the creation of multiple tasks
related to one problem for several teams to work in parallel. These
tasks are created by the team which leads the problem
investigation, they should include target completion time and
sufficient supporting information.
· In many cases, a temporary team combining different
competencies is created to start problem investigation. This is often
more effective than assigning parallel tasks to different teams,
especially if the problem is likely to be in an overlap of areas of
expertise and responsibilities. This technique is known as swarming.

Definition: Swarming

A technique for solving various complex tasks. In swarming, multiple people with different areas
of expertise work together on a task until it becomes clear which competencies are the most
relevant and needed.

During the problem control phase, one or more of the following is done, based on the
results of problem investigation:
· Previously unknown errors are identified, assessed, and assigned to
a relevant team for error control.
· Proactively identified errors are assessed and assigned to a relevant
team for error control.
· Recommendations for incident resolution based on an
understanding of the causes are developed, recorded, and
communicated to the relevant teams. These can be systemic
resolutions or workarounds for incidents.
· Problems with significantly low impact and probability and
problems that have been identified mistakenly are documented
and closed.

Definition: Workaround

A solution that reduces or eliminates the impact of an incident or problem for which a full
resolution is not yet available. Some workarounds reduce the likelihood of incidents.

Note that workarounds for incidents derived from problem analysis usually do
not reduce the likelihood of incidents. Instead, they help to resolve incidents
quicker and better when they occur. Workarounds that may help to prevent
incidents are more likely to be identified at the error control stage.
Problems that have not been dismissed at the problem control phase are
assigned the status of ‘known error’.

Definition: Known error

A problem that has been analysed but has not been resolved.

2.2.3 Error control

When a problem has been analysed (meaning the errors in the products have
been localized and their impact on services has been assessed), it should be
continually managed until resolved or closed without resolution.
Problem records may be closed only if one of the following conditions is met:
· The problem is solved: the risk of incidents associated with the
problem is removed or decreased to an acceptable level.
· The problem no longer affects the organization.

If an error in the organization’s products still has significant impact and


probability of related incidents, but does not have an acceptable resolution, it
should remain open and be regularly reviewed.
Note that although ‘known error’ is the state of a problem, some organizations
prefer to have separate records for problem control and error control. In these
cases, the problem record may be closed when problem analysis is complete,
and the following activities may be registered in a related known error record.
The above conditions for closure apply to known errors, whether it is a separate
record or a status of a problem record.
Many known errors remain open for a long time if they cannot be efficiently
resolved, and they keep affecting services. In these cases, the organization
may focus on maximizing the effectiveness and efficiency of incident handling
(sometimes to the level of fully automated detection and resolution), but the
problem records should remain open and periodically reviewed.
The above approach to error control is valid where the costs of problem
resolution may be higher than the costs of living with known errors and
effective incident management. This is typical for problems associated with
third-party components, especially where the third party is unresponsive, or
the components are no longer supported. Conversely, where components are
available for improvement and can be improved (especially software under the
organization’s own control), known errors should be quickly removed.
Known errors are a part of an organization’s technical debt and should be
removed, where reasonably practicable.

Definition: Technical debt

The total rework backlog accumulated by choosing workarounds instead of systemic solutions
that would take longer.

Error control ensures that the organization has sufficient up-to-date


information about all the known errors in its products, including their statuses
and their impacts on services.
Error control optimizes problem resolution so that its costs and side-effects
are balanced by its positive effects. Reviewing known errors periodically helps
to identify changes in circumstances (such as business impacts, the availability
of a permanent solution and the associated costs, and resource availability)
that may trigger the re-assessment of the error and initiate its resolution.
The key outputs of error control are improvement initiatives and change
requests, which initiate the resolution of problems. Some resolutions fix the
errors in configuration items and other product components. Others may
introduce permanent workarounds: changes to the product configuration
which do not fix the error, but reduce the likelihood of incidents to an
acceptable minimum. The associated problem records may then be closed,
but it is important to keep the knowledge about the errors available. This
knowledge may be extremely valuable for future service design and planning
of changes.
Permanent workarounds are normally used for components that the
organization cannot fix (legacy systems, engineering infrastructure provided
by third parties, and so on) but the use of permanent workarounds to prevent
incidents increases the organization’s technical debt and should be avoided
wherever possible.
To summarize, possible types of problem mitigation are listed in Table 2.2.
Table 2.2 Approaches to problem mitigation

Mitigation approach Applicability Effect

Full permanent fix of Recommended Incidents are prevented, side-


the errors found. approach for all CIs and effects are minimized, and the
other product quality of services is improved in
Problem record is components under the the short-, mid-, and long-term
closed. organization’s full perspectives.
control.

Highly recommended
for software developed
by the organization.

Permanent May be recommended Incidents are prevented for the


workaround isolating for CIs that cannot be current product configuration;
the errors. fixed (third-party and/or future designs and changes
legacy systems). should consider the workarounds
Problem record may and may be limited by them.
be closed or remain
open.
Solutions are Applicable to low- Incidents recur, but their impact
provided to optimize impact problems with is minimized. The known error
incident very high costs for should be periodically reviewed to
management. available permanent ensure that recommended
Problem record fixes or with no available incident solutions are effective
remains open. fixes. and there is still no permanent
problem solution available.

2.2.4 Problem models

Different sources and types of problem may require different approaches to


problem identification and control. For example, one or more of the following
problem types may require a special approach to the problem management
practice. These can be problems in:
· Information and technology
· Software
· Hardware

· data, including that which is sensitive


· Value streams and processes
· Procedures
· processes and ways of working

· Partners and suppliers


· third-party components
· contracts and agreements

· Organizations and people


· Skills and competencies

· Training

· The service consumer’s resources.


To optimize the handling and resolution of these and other types of problems,
service providers define problem models. Problem models help to manage
problems effectively and efficiently, often with better results because of the
application of relevant proven and tested methods.
Definition: Problem model

A repeatable approach to the management of a particular type of problem.

The creation and use of problem models are important activities in the problem
management practice. They are described in section 3.1.4.

2.3 Scope
The scope of the problem management practice includes:
· the identification and analysis of problems, including the analysis
and control of known errors
· the initiation of changes to fix or reduce the impact of problems
· the ownership and co-ordination of problem resolution

· providing information about problems to the relevant stakeholders

· monitoring of known errors and continual improvement of


workarounds.
There are several activities and areas of responsibility that are not included in
the problem management practice, although they are still closely related to
problems. These are listed in Table 2.3, along with references to the practice
guides in which they can be found.
Table 2.3 Activities related to the problem management practice described
in other practice guides

Activity Practice guide

Incident resolution Incident management


Control and implementation of changes initiated to Change enablement
fix the problems
Deployment management

Infrastructure and platform


management

Release management

Software development and


management

Other practices

Risk assessment and control Risk management

Detection and control of errors in products before Deployment management


deployment to the live environment
Service design

Service validation and


testing

Software development and


management

Communication of workarounds for incidents to users Service desk

2.4 Practice success factors

Definition: Practice success factors

A complex functional component of a practice that is required for the practice to fulfil its
purpose.
A practice success factor (PSF) is more than a task or activity, as it includes
components of all four dimensions of service management. The nature of the
activities and resources of PSFs within a practice may differ, but together they
ensure that the practice is effective.
The problem management practice includes the following PSFs:
· identifying and understanding the problems and their impact on
services
· optimizing problem resolution and mitigation.

2.4.1 Identifying and understanding the problems and their impact on services

Organizations should understand the errors in their products because they


may cause incidents and affect service quality and customer satisfaction. The
problem management practice ensures problem identification and thus
contributes to the continual improvement of products and services. This is
more effective if performed proactively rather than reactively.
Organizations starting out with problem management usually begin with
reactive problem identification based on repeating and major incidents. The
next steps in capability development include proactive problem identification
and expanding the scope of problem management beyond the information
and technology dimension. More information on the problem management
capability levels can be found in section 7.3.
It is important to find the right balance between the accessibility and
effectiveness of problem identification. Although problems can, and should,
be identified in many areas of the service provider’s organization (by
developers, support teams, technical specialists and others) it is important to
ensure that there is accurate and sufficient information about identified
problems, and that they are assigned correctly.
This balance can be achieved by appropriate training, or by assigning
responsibility for problem identification to people in each relevant team.

Key message

Who can register a problem?


There are several approaches to assigning responsibility
There are several approaches to assigning responsibility
for problem registration. One approach is to encourage
every practitioner, analyst or specialist to initiate and
register problems. This would increase the number of
improvements and improve the visibility of the errors in
the organization’s products.

However, this approach can be limiting and frustrating


where an individual wants to raise a problem but is
unable to easily do so. This may also significantly increase
the number of registered problems that are not actively
worked on, or that are incorrectly categorized. To prevent
this, some organizations prefer to make one or more roles
responsible for the initial filtering and registering of
potential problems. This approach may be effective as
long as the individuals carrying out these roles have
sufficient resources and authority, and can process
information from various sources consistently and
transparently.

Organizations can combine different approaches to


balance the scope, throughput, and efficiency of problem
identification.

It should be accepted that in some cases problems will be incorrectly


identified, categorized, or assigned. Blame should not be apportioned for
these mistakes, but instead this should lead to analysis and improvement of
the problem models and problem identification guidelines.
Teams should understand the value of problem management for their
product or technology domain, and eventually for the organization. Problem
management should be treated as an important form of risk management
and of continual improvement. If these practices are formalized in the
organization, problem management should follow their recommendations
and be aligned with them in the context of the organization’s service value
system.
2.4.2 Optimizing problem resolution and mitigation
When problems have been identified, they should be handled effectively and
efficiently. It is rarely possible to fix (remove) all problems in all organization’s
products, but identification without resolution is significantly less valuable for
the organization and its customers. A balanced approach should be defined
for problem mitigation, considering the associated costs, risks, and impacts on
the service quality. It is important to make decisions about problem resolution
or mitigation based on the business impact of different scenarios, rather than
purely technical considerations.
As problem investigation and resolution may require significant resources, it is
important to maintain awareness and support of the practice at all levels of
the service provider.
Some organizations find it useful to have and share across the teams a list of
the most important open problems. This is a simple yet effective way to:
· demonstrate management attention to the problems
· raise awareness of the prioritized problems across the service
provider
· involve people from different teams in suggesting ideas and
solutions
· keep track of the status and progress of the problem investigation
and resolution
· keep business stakeholders informed and maintain their interest
and support.

2.5 Key metrics


Key metrics for the problem management practice are mapped to its PSFs.
The key metrics are listed in Table 2.4.
The practice metrics should be applied to a specific context such as type of
problems, services, technical domain, or periods of time.
The effectiveness and performance of the ITIL practices should be assessed
within the context of the value streams to which the practices contribute. The
context of the business and the value streams is important when defining
whether the practice’s performance is considered good or not. This is why this
practice guide cannot recommend universal key performance indicators for
problem management: the target values for each metric can only be defined
in the organization’s context.
Table 2.4 Key metrics of problem management
Practice success factors Key metrics

Identifying and understanding the Number and impact of problems identified


problems and their impact on services over the period

Number and impact of incidents that are


not associated with known errors

Number and impact of incidents that


require urgent problem investigation

Optimizing problem resolution and Number and impact of incidents prevented


mitigation by problem resolution

Number and impact of incidents resolved


with solutions provided by problem
investigation

Number and impact of known errors that


remain open

The correct aggregation of metrics into complex indicators will make it easier
to use the data for the ongoing management of value streams, and for the
periodic assessment and continual improvement of the problem
management practice. There is no single best solution. Metrics will be based
on the overall service strategy and priorities of an organization, as well as on
the goals of the value streams to which the practice contributes.
Note: A useful metric for problem management was proposed by D. Isaychenko and P. Demin[1].
They called it a problem management performance index (PPI):
PPI = N+R ÷ O+C [0;1] ↑

Where
R is the total number of problems resolved during the reporting period;
O is the total number of problems that were open at the end of the period;
N is the number of problems logged during the period and still open by the end of the period;
C is the total number of problems closed during the period. C includes actually resolved
problems and problems that were closed without being resolved. . To avoid accidental
distortions, both R and C do not include problems logged by mistake, such as duplicates or
unintentional logging, which are those problems that were closed without any actual
processing.
PPI increases with new logged problems as well as with resolved ones. This demonstrates and
can be used to stimulate both problem identification and resolution.

[1][Link]

3. Value streams and processes


3.1 Processes
Each practice may include one or more processes and activities that may be necessary
to fulfil the purpose of that practice.

Definition: Process

A set of interrelated or interacting activities that transform inputs into outputs. A process takes
one or more defined inputs and turns them into defined outputs. Processes define the sequence
of actions and their dependencies.

Problem management activities form four processes:


· proactive problem identification
· reactive problem identification

· problem control
· error control.

3.1.1 Proactive problem identification


This process includes the activities listed in Table 3.1 and transforms the inputs into
outputs.
Table 3.1 Inputs, activities, and outputs of the proactive problem
identification process

Key inputs Activities Key outputs

Error information from vendors and Review of the Problem


suppliers Information about potential submitted information records
errors submitted by specialist teams
Problem registration Feedback to the
Information about potential errors problem
submitted by external user and Initial problem initiator
professional communities categorization and
assignment
Information about potential errors
submitted by users Monitoring data

Service configuration data

Figure 3.1 shows a workflow diagram of the process.


Figure 3.1 Workflow of the proactive problem identification process

Proactive problem identification is used to identify potential errors in the


organization’s products based on sources other than incident records.
Proactive problem identification and control can be considered and
performed as a form of risk management which is focused on the
vulnerabilities in the organization’s product: it includes the identification,
assessment, and analysis of the vulnerabilities and the associated risks.
Possible sources of information about errors in an organization’s products are
listed in Table 3.2.
Table 3.2 Sources of information for proactive problem identification

Source Examples of information


Service designers, software developers, Errors in the current live versions
architects, and other teams working on discovered during work on the subsequent
the next versions of CIs and other versions
components
Errors in the versions currently being
deployed to the live environment that
have been identified during testing but
have not been fixed

Vendors of software and other CIs Errors in the current live versions of the
vendor’s systems and components

User and professional communities Errors by other organizations using the


same versions of systems and components

Monitoring data Suspicious trends and deviations in the


performance of services and CIs

Users Vulnerabilities in the services being used

Where possible, proactive problem identification should focus on key systems


and components which have the highest potential impact on the organization
and its customers.
However, indications of errors in other systems should not be neglected. In
complex technical environments designed for high availability, incidents may
have multiple causes which are often the result of improbable combinations
of improbable factors. Seemingly unimportant errors in non-core systems can
contribute to serious incidents. Proactive problem identification should
include the careful assessment of the probability and impact of the identified
vulnerabilities. Table 3.3 provides a description of the process activities.
Table 3.3 Activities of the proactive problem identification process

Activity Description
Review of the Depending on the source and the subject, the submitted
submitted information is reviewed by a specialist or a specialist group. The
information review includes checks for duplicates, applicability, common
sense, and ongoing incidents potentially related to the
submitted information.
If the decision is made not to register a problem, the initiator
may be notified (usually applicable in case of an active or ‘push’
submission; not applicable if the information was obtained or
‘pulled’ from external sources, such as vendor bulletins where
nobody is expecting feedback).

Problem If the need for problem control is confirmed, a problem record is


registration registered. This can be done by a dedicated role or by a wider
group of specialist roles.

Initial problem The person registering a problem performs the initial


categorization and categorization. The information usually includes some of the
assignment following (if known or reasonably assumed):
· source
· description
· associated CIs and/or CI classes
· estimated impact and probability of incidents
· associated and potentially affected services
· impact on the organization and customers.

Based on the initial categorization, the problem is assigned to a


specialist group responsible for the associated CI, service, or
product. Where applicable and expected, the problem initiator
may be notified about the problem registration.

3.1.2 Reactive problem identification

This process includes the activities listed in Table 3.4, and transforms the inputs into
outputs.
Table 3.4 Inputs, activities, and outputs of the reactive problem
identification process
Key inputs Activities Key outputs

Information about ongoing Problem registration Problem


incidents records
Initial problem categorization and
Incident records and reports assignment

Monitoring data

Service configuration data

Service level agreements


(SLAs)

Figure 3.2 shows a workflow diagram of the process.

Figure 3.2 Workflow of the reactive problem identification process

Reactive problem identification uses information about past and ongoing


incidents to investigate their causes. It can be triggered by an ongoing
incident investigation that could not identify the nature of the incident; in this
case, problem identification and control may be urgent. The incident
management and problem management practices are used within a single
value stream and are likely to involve the same (or overlapping) resources,
including teams, tools, and procedures.
When based on the analysis of past incidents, problem identification may
include statistical analysis, impact analysis, and trend analysis from various
perspectives. The aim is to identify groups of incidents with possible common
causes and significant total business impact.
This process varies slightly depending on the trigger. The variations are
illustrated in Table 3.5.
Table 3.5 Activities of the reactive problem identification process

Activity Triggered by ongoing incident Triggered by incident records


analysis
Problem The team working on the incident A specialist team responsible
registration identifies the need for problem for a system, service, or
investigation. In some cases, a product performs regular
problem record is linked to one or reviews of the incident records
more incident records for tracking associated with their area of
the investigation. It may be responsibility. If they detect a
especially important where reason for a problem
multiple incidents in numerous investigation, they register a
locations are being handled by problem record. These reasons
different teams and require a may include:
coordinated problem · a high number
investigation, or where problem of similar
investigation will be done by a incidents
dedicated team. · a high
percentage of
In other cases, the team working incidents
with the incident may continue resolved after
investigating the incident’s causes the target
and document the problem after resolution time
the incident is resolved. The · major incidents
problem may still need to be · availability below
registered, especially if the causes the target level.
of the incident were not removed
during incident resolution and
new incidents may arise from the
same problem.
Initial problem When registering a problem, the When registering a problem,
categorization person doing so performs initial the person doing so performs
and categorization. This usually initial categorization. This
assignment includes some of the following (if usually includes some of the
known or reasonably assumed): following (if known or
· description reasonably assumed):
· associated CIs · description
and/or CI classes · associated
· estimated impact incidents and
and probability of their solutions
incidents · associated CIs
· associated and and/or CI classes
potentially affected · estimated
services impact and
· impact on the probability of
organization and future incidents
customers · associated and
potentially
If the problem is registered before affected services
the problem investigation, the · impact on the
problem is assigned to the organization
appropriate specialist group. and customers
· estimated
If the problem is registered after impact and
the problem investigation, the probability of
information includes the steps incidents
made, the results, and the current
status of the problem. If the Based on initial categorization,
problem is not solved at the time the problem is assigned to a
of registration, it is assigned to the specialist group, responsible
appropriate group. for the associated CI, service, or
product.

3.1.3 Problem control

This process focuses on the investigation of the problem. It includes the activities shown
in Table 3.6 and transforms the inputs into outputs.
Table 3.6 Inputs, activities, and outputs of the problem control process
Key inputs Activities Key outputs

Problem records Problem investigation Problem


records
Service configuration data Known error
communication Known errors
Technical information about CIs,
products, and services Incident
solutions
Incident records

Monitoring data

Figure 3.3 shows a workflow diagram of the problem control process.

Figure 3.3 Workflow of the problem control process

Table 3.7 Activities of the problem control process


Activity Description

Problem The specialist team assigned to the problem investigates the


investigation possible causes of the incident and/or verifies the reported errors
in the CIs and the organization’s other resources. . The methods
and procedures vary depending on how the problem has been
identified. For problems identified reactively, localization starts
with understanding which CIs may have errors causing past or
ongoing incidents. For most problems identified proactively, CIs
or CI classes would have been identified during their registration.

After the problem is localized to the level of CIs, further


diagnostics may be needed to identify errors within the
suspicious CIs. This and the following activities may be performed
by different teams (teams re-assigned based on the problem
localization). If the reported problem is irrelevant to the
organization (for example, a publicly reported vulnerability in
software that does not affect the versions used by the
organization), the problem record may be closed.

If the investigated problem is relevant to the organization, it is


assigned the known error status for further control and
resolution. Actions and results of the investigations are recorded
in the problem records.

Known error The results of problem investigation are communicated to the


communication problem initiator and relevant teams. These may include product
development teams, technical specialists, the service desk team,
users, and suppliers.

If there are ongoing incidents associated with the problem that is


being investigated, the results of the problem localization are
communicated to the incident investigation teams.

It is possible that understanding the errors is enough to define an


incident resolution. In this case, a recommended solution for the
incident should be registered in the problem records and
communicated to the teams working with the incident.
To investigate problems, organizations use various analysis techniques. These may
include:
· root cause analysis techniques, such as 5 Whys, Kepner and Fourie,
and fault tree analysis
· impact analysis techniques, such as component failure impact
analysis and business impact analysis
· risk analysis techniques.
It is important to remember that the concept of a single root cause has a very
limited applicability in complex evolving environments. Quite often, incidents
are caused by improbable combinations of improbable factors. This means
that investigation of problems (especially identified reactively) should not be
limited to the identification of the first possible cause of incidents. Problem
investigation should always consider all four dimensions of service
management.
3.1.4 Error control

This process focuses on the monitoring and control of the status of the known
errors (problems that are analysed but not resolved) and their resolution. It
helps to ensure that the negative impacts of the known errors on services are
understood and minimized; the solutions for related incidents are effective;
and the mitigation approach for the known error is valid, effective, and
efficient.
This process includes the activities shown in Table 3.8 and transforms the
inputs into outputs.
Table 3.8 Inputs, activities, and outputs of the error control process

Key inputs Activities Key outputs


Problem records Problem solution Problem records
development
Service configuration data Problem models
Problem resolution
Technical information about CIs, initiation Change requests
products, and services
Known error monitoring Improvement
Incident records and review initiatives

Monitoring data Problem closure Problem solutions

Knowledge management data

Figure 3.4 shows a workflow diagram of the process.

Figure 3.4 Workflow of the error control process

Table 3.9 Activities of the error control process

Activity Example
Problem The team (assigned or re-assigned based on the problem
solution investigation) looks for a way to solve the problem. This includes
development defining an approach to the mitigation (see Table 2.1) and
development of the actual solution within the selected approach. If
there is no viable solution for the problem, the supporting
information is recorded and the error goes to periodic review.

Problem In most cases, problem resolution requires change. The responsible


resolution team submits change requests, following the organization’s
initiation procedures for change initiation and implementation.

In other cases, required actions are not classified as changes and


can be initiated and performed following other procedures. Either
way, the team initiates the actions required for the defined (and, if
needed, approved) problem resolution. This initiation may need to
be supported with relevant justification (including financial, risk,
compliance, technical, and other considerations).
Known error If a solution is approved for the known error
monitoring and The implementation of the solution is controlled and confirmed
review using pre-agreed criteria. This is usually done by the team that
initiated the resolution, or another pre-agreed role, such as
problem manager.

For reactively identified problems, this can be done based on the


change in incident dynamics (related incidents are resolved or
prevented). For proactively identified problems, resolution control is
based on the success of the initiated changes and may include a
period of monitoring any service that might have been affected by
the errors.

If the resolution of the problem is unconfirmed, the team returns to


the problem solution development step of the process.

If no viable solution is found for the known error


An assigned specialist team should monitor the known error. This is
usually the team responsible for the CI, service, or product with
which the known error is associated. The team monitors the status
of the known error as defined in the mitigation strategy. Monitored
parameters may include:
· the dynamics of the associated incidents
· the effectiveness of the incident solutions
· the effectiveness of the problem workarounds
· changes in the statuses of the resources needed
to solve a problem (budget, updates from the
vendor, specialists, new infrastructure, and so
on).

The team should conduct problem reviews periodically (in line with
the agreed mitigation approach) or based on outstanding
monitoring results.

If the review confirms that the mitigation approach is valid and up


to date (the problem exists, the impact assessment is up to date,
incident solutions are effective, the problem workaround is
effective, and no viable problem fix is available), then known error
monitoring continues.
If the mitigation approach becomes invalid, the problem solution
development activity is initiated to review and redefine the
mitigation approach. This may include developing and
implementing a problem solution or updating the incident
solutions for any associated incidents.

If the problem no longer exists (for example, it has been removed


with planned software or hardware updates or by
decommissioning the affected CIs), problem closure is initiated.

If the problem demonstrated a new pattern that suggests the


amendment or creation of a problem model, a problem model is
documented and communicated as part of the problem review
activity.

Problem records are updated with monitoring data.

Problem The specialist (or team) acting as a problem manager/coordinator,


closure reviews the results and formally closes the problem record.

If the resolution is confirmed, they document the resolution control


results and formally close the problem record.

Closed problem records should be available as part of the


organization’s knowledge base, especially if there is a chance that
similar problems may recur.

3.2 Value stream contribution


3.2.1 Service value streams

To perform certain tasks or respond to particular situations, organizations


create service value streams. These are specific combinations of activities and
practices, and each one is designed for a particular scenario. Once designed,
value streams should be subject to continual improvement.

Definition: Value stream


A series of steps an organization undertakes to create and deliver products and services to
consumers.

In practice, however, many organizations come to use of the value stream


concept after having worked for a while (sometimes for years) without the
value streams being managed, mapped, or understood. This means that when
the importance of the concept becomes clear, the first step is to understand
and map the ‘As Is’ situation, the de-facto flows of work, and to analyse them
in order to identify and eliminate the non-value-adding activities and other
forms of waste.
Identifying and understanding the existing value streams is critical to
improving organization’s performance. Structuring the organization’s activities
in the form of value streams allows it to have a clear picture of what it delivers
and how, and to make continual improvements to its services.
Combined, organizations’ value streams form an operating model which can
be used to understand and improve how the organization creates value for the
stakeholders.
Many organizations follow best practice recommendations for various service
management practices, such as incident management, change enablement,
software development, and many others. Problem management is one of the
most adopted and developed practices.
However, the practices have often been adopted and organized in a siloed,
isolated manner, just as they were presented in the service management
bodies of knowledge. In reality, a flow of work required to create or restore
value for a customer or another stakeholder is almost never limited to one
practice.
The project may be part of a programme or portfolio structure, or it may be a
standalone project reporting to the business unit’s management structure, as
illustrated in Figure 1.2.
3.2.2 Problem management in service value streams

Problem management is likely to be involved in several service value streams.


The most obvious one is the restoration of normal operations in case of an
incident. This service value stream is not limited to the incident management
practice, but involves many other practices, including problem management.
Table 3.10 Practices in the service value stream restoring normal operation
after an incident
Activity Practice

Incident detection Service desk (for user-reported incidents)


or
Monitoring and event management

Incident registration Incident management

Incident classification

Incident diagnosis Incident management

Knowledge management

Problem management

Incident resolution Incident management and one or more of:


Problem management

Change enablement

Software development and management

Service validation and testing

Deployment management

Release management

Service desk

Infrastructure and platform management

Supplier management
Incident closure Incident management

Service desk

Monitoring and event management

Problem management

Knowledge management

Business relationship management

In this value stream, the problem management practice may be used to:
· provide information about currently open problems
· provide information about previously identified solutions for similar
incidents
· urgently investigate the causes of the incident(s) and find ways to
resolve the incident(s)
· capture information about the incident(s) to update knowledge
about errors, symptoms, and resolutions.
Similarly, problem management may be involved in other value streams,
including:
· design, development and implementation of new and changed
products and services
· routine operations of IT systems.
In all cases, the practice is involved for the same purposes listed above:
providing previously identified solutions and other relevant information or
capturing information about new problems.
There is no single operating model fitting all organizations. Different solutions
work for different organizations, involving different value streams which in
turn involve different management practices.
3.2.3 Analysing a service value stream

[Link] The key steps of a service value stream analysis


The following are some simple and practical recommendations for service
value stream analysis and mapping.
1. Identify the scope of the value stream analysis: Value streams can
be mapped to a particular product or service or applied to most or
all of them. Value streams may differ for different consumers; for
example, incidents can be solved and communicated differently for
internal and external customers, or for B2B and B2C products, or for
services based on products developed inhouse or sourced
externally.
2. Define the purpose of the value stream from the business
standpoint: Make sure the stakeholder’s concerns are clearly
understood, since they are the ones defining value. In the case of
service desk, it is usually users who need a convenient interface to
communicate with the service provider; however, there are usually
other interested parties. For example, internal users may be unable
to provide normal service to a business customer because of the
incident, and the value of the value stream should be considered
from the business perspective, not solely from the user perspective.
3. Do the service value stream walk: Walk or directly experience the
steps and information flow as they go in practice (consider the Lean
technique of Gemba walk):
a. Identify the workflow steps
b. Collect data as you walk
c. Evaluate the workflow steps: Typically, the criteria for evaluation
are:
· value for the stakeholder (does the step add value for the
business stakeholder?)
· effectiveness and performance (is the step performed well?)
· availability (are required resources available to execute the
step?)
· capacity (are required resources enough?)
· flexibility (are the required resources interchangeable within
the step?).
d. Map the activities and the information flows: In an ideal
situation, the flow goes smoothly without delays and pauses,
there are no disconnections between the steps, and the
workload is level with minimal (and agreed) variation.
e. Create and review the timeline and resource level: Map out
process times and lead times for resources and workload
through the workflow steps.
4. Reflect on the value stream map (VSM): Identify factors that
might not have been entirely apparent at first. The information
collected is used at this step to find the waste.
5. Create a ‘to be’ VSM: This informs and drives improvement. The
value stream should be considered holistically to ensure end-to-
end efficiency and value creation, not just local improvements.
6. Using the ‘to be’ VSM, plan improvements: Refer to the continual
improvement practice guide for a practical improvement model.
[Link] Problem management considerations in a service value stream
analysis
To ensure that relevant problem management activities are included in
service value streams, the following steps can be added to the above
recommendations.
· At the scoping step (1), identify the IT and business services related
to the value stream and the involved business stakeholders. For
example, should the problem management practice be involved in
the investigation of problems causing business service incidents, or
should it be limited to problems in IT services?
· Make sure the value stream is understood (step 2) from the
standpoint of the business, not only of the service provider.
· During the service value stream walk (3a), identify other practices
involved in dealing with problems at every step. Which practices
provide required information (configuration data, asset data,
financial information, agreed timelines for the service restoration…)?
How are changes initiated if required for problem resolution? What
if a resolution initiates a project? How are third parties involved in
problem investigation and error control? How information about
vulnerabilities in third- party products is obtained from the
suppliers?
· During the workflow steps evaluation (3c), evaluate the step’s
impact on the business value. Special attention should be paid to
steps with low business value, low performance, and availability or
capacity issues. It is not unusual to find steps which serve some
internal control or bureaucratic purposes but delay the incident
resolution.
· At the reflection and planning steps (4-5), ensure that the problem
management processes are optimized for business value
throughout the stream, not only at the within the problem
management practice.
· Consider including the creation or updating of problem models
(see section 2.2.4) in the value stream improvement plans (step 6).

4. Organizations and people


4.1 Roles, competencies, and responsibilities
The practice guides do not describe the practice management roles such as
practice owner, practice lead, or practice coach. They focus instead on the
specialist roles that are specific to each practice. The structure and naming of
each role may differ from organization to organization, so any roles defined in
ITIL should not be treated as mandatory, or even recommended.
Remember, roles are not job titles. One person can take on multiple roles and
one role can be assigned to multiple people.
Roles are described in the context of processes and activities. Each role is
characterized with a competency profile based on the model shown in Table
4.1.
Table 4.1 Competency codes and profiles

Competency Competency profile (activities and skills)


code

L Leader: Decision-making, delegating, overseeing other activities,


providing incentives and motivation, and evaluating outcomes

A Administrator: Assigning and prioritizing tasks, record-keeping,


ongoing reporting, and initiating basic improvements
C Coordinator/communicator: Coordinating multiple parties,
maintaining communication between stakeholders, and running
awareness campaigns

M Methods and techniques expert: Designing and implementing


work techniques, documenting procedures, consulting on processes,
work analysis, and continual improvement

T Technical expert: Providing technical (subject matter) expertise and


conducting expertise-based assignments

Two practice-specific roles may be found in organizations: problem manager


and problem coordinator.
These roles are often introduced in organizations where the number of
problems is high. In other organizations, problem management activities are
coordinated by a person or a team responsible for the CIs, service, or product
with which the problem is associated; this may be the resource owner, service
owner, or product owner respectively.
4.1.1 Problem manager role

Where a dedicated problem manager role is defined, it is usually assigned to


specialists combining good knowledge of the organization’s products
(architecture, configurations, and interdependencies) with solid analytical and
leadership skills (the ability and authority to coordinate teamwork and provide
good risk management). The competency profile for this role is CLMT.
This role is usually responsible for managing and coordinating the specialist
activities in the problem management processes, including:
· conducting and coordinating problem registration based on the
submitted information
· the initial categorization of the problems
· coordinating problem investigation and solution implementation
control
· coordinating the communication with the teams responsible for
incident resolution and change implementation
· developing and communicating problem models, where applicable
· coordinating known error monitoring and review
· the formal problem closure.

4.1.2 Problem coordinator role

In more complex organizations, some responsibilities for the problem


management practice may be delegated to the problem coordinator. The
problem coordinator focuses on routine problem management activities, such
as the review of submitted information about possible problems, problem
review, and problem closure.
Examples of other roles which can be involved in the problem management
activities are listed in Table 4.2, together with the associated competency
profiles and specific skills.
Table 4.2 Examples of roles with responsibility for problem management
activities

Activity Responsible Competency Specific skills


roles profile

Proactive problem identification process

Review of the CI owner T Good knowledge of the


submitted product, including its
information Problem architecture and configuration
coordinator

Problem
manager

Product
owner

Service
owner
Problem CI owner TA Knowledge of the registration
registration tools and procedures
Problem
coordinator

Problem
manager

Product
owner

Service
owner

Initial problem CI owner TAC Good knowledge of the


categorization and product, service architecture,
assignment Problem and business impact
coordinator Understanding of the
responsibilities and
Problem competencies across the teams
manager

Product
owner

Service
owner

Reactive problem identification process


Problem CI owner TA Knowledge of the registration
registration tools and procedures
Incident
manager

Problem
coordinator

Problem
manager

Product
owner

Service
owner

Initial problem CI owner TAC Good knowledge of the


categorization and product, service
assignment Incident architecture, and business
manager impact Understanding of
the responsibilities and
Problem competencies across the
coordinator teams

Problem
manager

Product
owner

Service
owner

Problem control process


Problem CI owner CT Good knowledge of the
investigation product, service architecture,
Problem and business impact.
coordinator
Good knowledge of diagnostic,
Problem investigation, and analysis
manager methods and tools.

Product
owner

Service
owner

Supplier

Technical
specialist

Known error CI owner TC Understanding of stakeholders


communication and responsibilities
Incident
manager Knowledge of the
communication tools and
Problem procedures
coordinator

Problem
manager

Error control process


Problem solution CI owner TMC Good knowledge of the
development product and service
Problem architecture, configuration, and
coordinator technical details

Problem Creativity
manager
Systems thinking
Product
owner

Service
owner

Supplier

Technical
specialist

Problem CI owner CT No specific skills required


resolution
initiation Problem
coordinator

Problem
manager

Product
owner

Service
owner

Supplier

Technical
specialist
Known error CI owner TAC Good knowledge of the
monitoring and product and service
review Problem architecture and business
coordinator impact

Problem
manager

Product
owner

Service
owner

Supplier

Technical
specialist

Problem closure CI owner TCA Good knowledge of the


product, service architecture,
Problem and business impact
coordinator

Problem
manager

Product
owner

Service
owner

4.2 Organizational structures and teams


It is unusual to see a dedicated organizational structure for the problem
management practice, although the role of problem manager is sometimes
associated with a formal job title. This is typical for organizations with complex
bureaucracy and a significant number of problems to manage. Many
organizations find it useful to form temporary teams to investigate high-
impact problems and/or to develop solutions.
In product-focused organizations, problem management job titles and roles
are not typically adopted. Instead, this practice is integrated in the day-to-day
activities of the product development and management teams. It is
automated wherever possible.
Problem management is a cross-functional collaborative practice, and
requires support and participation from a range of teams and people across
an organization. As such, if there is a specialized position or team fully
dedicated to problem management, they become responsible for leading
problem management activities and involving other teams in the practice.
Problem management relies on information from, and the participation of, all
teams within a service provider, and often also of its suppliers and customers.
An example of a coordinating activity that may be performed by the dedicated
problem manager is organizing a swarm (a temporary team assembled on a
short notice to investigate a problem).

5. Information and technology


5.1 Information exchange
The effectiveness of the problem management practice is based on the quality of the
information used. This information includes, but is not limited to, information about:
· products and services and their architecture and design, including
configuration information
· customers and users
· partners and suppliers, including contract and SLA information on
the services they provide
· ongoing and past incidents
· planned, ongoing, and past changes
· a third-party’s products and components, including vulnerabilities
and incidents.
This information may take various forms. The key inputs and outputs of the practice are
listed in chapter 3.
5.2 Automation and tooling
The problem management practice can significantly benefit from automation. The term
automation is used in this and other ITIL publications to refer to the use of digital
technology to enable, support, or enhance various activities. This includes, but is not
limited to the full automation of activities where technology solutions remove the need
for human intervention. Table 5.1 provides a list of the key automation supporting the
practice and their most common application.
Table 5.1 Automation solutions for the problem management practice

Automation tools Application in problem management

Workflow management Reactive problem identification (analysis of incident


and collaboration tools records)

Management of problem and known error records

Support of collaboration during problem investigation


and resolution

Support of problem impact analysis and root cause


analysis (through records of other management practices,
including incidents and changes)

Monitoring and event Proactive problem identification


management tools
Support of trend analysis for problem impact assessment

Monitoring of the effectiveness of the resolution

Service configuration Problem categorization and investigation


management tools

Analysis and reporting Problem investigation


tools Practice measurement and reporting
Knowledge Retrieving, managing, and communicating known
management tools solutions for incidents and problems

Detailed descriptions of how these tools support the practice’s activities are outlined in
Table 5.2.
Table 5.2 Details of automation of the problem management activities

Process activity Means of Key functionality Impact on the


automation effectiveness of
the practice

Proactive problem identification process

Review of the Monitoring and Collection and overview of High


submitted event information from various
information management sources, including data
tools analysis and team
collaboration
Workflow
management
and
collaboration
tools

Problem Workflow Management of problem High


registration management records integrated with
and other service
collaboration management data
tools
Initial problem Workflow Management of problem High
categorization management records integrated with
and assignment and other service
collaboration management data,
tools backlog management,
communication, and
Service collaboration support
configuration
management
tools

Reactive problem identification process

Problem Workflow Machine-learning-based High


registration management problem identification
and based on analysis of past
collaboration and ongoing incidents
tools
Management of problem
records integrated with
other service
management data

Initial problem Workflow Management of problem High


categorization management records integrated with
and assignment and other service
collaboration management data,
tools backlog management,
communication,
Service collaboration support, and
configuration CI impact assessment
management
tools

Problem control process


Problem Analysis and Dependencies analysis, High
investigation reporting tools what-if analysis, cause
and-effect analysis, and
Service modelling
configuration
management
tools

Known error Workflow Communication and Medium


communication management collaboration support
and
collaboration
tools

Error control process

Problem solution Analysis and Solution design and Medium to very


development reporting tools validation high, depending
on the solution
Service architecture
configuration
management
tools

Problem Workflow Communication and Medium


resolution management collaboration support
initiation and
collaboration
tools
Known error Monitoring and Collection and overview of Medium to high
monitoring and event information from various
review management sources, data analysis, and
tools team collaboration

Workflow Verification that known


management errors exist and
and workarounds work
collaboration
tools Creating and linking
knowledge articles to
known errors

Problem closure Workflow Communication and Medium


management collaboration support,
and automatic posts into
collaboration collaboration tools
tools

5.2.1 Recommendations on automation of problem management

The following recommendations can help when applying automation to


problem management:
· Distinguish between problem control and error control: Whether
problems and known errors are different objects or different
statuses of the same object in the automation tool, make sure they
support clear separation between investigation (problem control)
and periodic review and resolution (error control).
· Ensure integration with other practices: At the least, it should be
possible to link problem records to incident records, change
records, configuration items, and services.
· Ensure integration with knowledge base(s): At the least,
information about errors and related recommendations (symptoms
of possible incidents, workarounds) should be available. Consider
integration (and automated analysis) of the vendors’ and suppliers’
knowledge bases, bug reports, release notes and so on. If
applicable, capture information shared by other users of the same
third-party products about their experience, problems and
resolutions.
· Pay attention to measurement and reporting from the
beginning: Ensure visibility of the problem management status
and progress, including the four processes (reactive and proactive
identification, problem control and error control).
· Leverage machine learning capabilities: Problem identification
based on incident records or on monitoring data; cause-effect
analysis; impact analysis, search for resolutions, and other activities
of the problem management practice can be automated with the
use of machine learning (ML). Effective use of ML requires high-
quality data and effective integration with various sources of
information. If used properly, it can significantly improve the
problem management practice.
· Automate monitoring of known errors: For known errors requiring
periodic control, the conditions of the control can be defined for
automated monitoring. For example, if the error is expected to be
resolved after a software update is released by a vendor, it is
possible to automate monitoring of the updates and notification (or
assignment) to a relevant specialist. If a resolution has been applied
and a certain type of incident is supposed to be prevented by the
resolution, incident monitoring can automatically reassign the
problem for further investigation (if the incidents occur), or for
review and closure (if the incidents do not occur for an agreed
period of time).

6. Partners and suppliers


Very few services are delivered using only an organization’s own resources.
Most, if not all, depend on other services, often provided by third parties
outside the organization (see section 2.4 of ITIL® Foundation: ITIL 4 Edition for
a model of a service relationship). Relationships and dependencies introduced
by supporting services are described in the ITIL practices for service design,
architecture management, and supplier management. Information about
dependencies on third-party services is used in all problem management
processes.
Partners and suppliers may support the development, management, and
execution of the problem management practice. The forms of support include
the following:
· Performing problem management activities: Some problem
management activities can be largely or completely performed by
a specialized supplier. Third parties are often involved in problem
investigation and resolution. It is important to ensure effective
integration of these third parties in the problem-related workflows
and information exchange, as well as their adherence to relevant
policies.
Problem models should define how third parties are involved in
problem control and how the organization ensures effective
collaboration with them. This depends on the architecture and
design solutions for products, services, and value streams. Quite
often, after the correct model is selected for a problem, further
consideration of third-party dependencies is needed within the
processes of problem and error control.
The problem management practice often discovers errors in third-
party products used by the organization. The possibility of solving
these errors, and the effectiveness of the solution, depends on
multiple factors, including:
· the architecture of the solution
· the flexibility of the supplier
· the importance of the service relationship with the organization
for the supplier
· the contract terms and conditions.
It is important to understand how the organization depends on third-party
components and how it aims to establish effective and efficient collaboration
with its key suppliers and partners around many activities, including those of
the problem management practice.
Where organizations aim to ensure fast and effective problem management,
they usually try to agree close cooperation with their partners and suppliers,
removing formal bureaucratic barriers in communication, collaboration, and
decision-making (see the supplier management practice guide for more
information).
· Provision of software tools: Most software tools used for problem
management are shared with other practices. Problem
management is never the first management practice to be
automated and rarely one of the most important; this sometimes
results in problem management requirements being deprioritized,
and the practice affected by ineffective or partial use of software
tools. It is important to ensure that specific requirements such as
the aggregation of problem information from different sources,
cross-team collaboration, or changeable impact and categorization,
are met by the software tools.
· Consulting and advisory: Specialized suppliers who have
developed expertise in problem management can help to establish
and develop the practice, adopt methods and techniques (such as
swarming or ‘5 Why’ analysis), and initially develop the problem
models.

7. Capability assessment and development


7.1 The practice capability levels
The practice success factors described in section 2.4 cannot be developed
overnight. The ITIL maturity model defines the following capability levels
applicable to any management practice:
Level 1 The practice is not well organized; it’s performed as initial or intuitive. It
may occasionally or partially achieve its purpose through an incomplete set of
activities.
Level 2 The practice systematically achieves its purpose through a basic set of
activities supported by specialized resources.
Level 3 The practice is well defined and achieves its purpose in an organized
way, using dedicated resources and relying on inputs from other practices that
are integrated into a service management system.
Level 4 The practice achieves its purpose in a highly organized way, and its
performance is continually measured and assessed in the context of the
service management system.
Level 5 The practice is continually improving organizational capabilities
associated with its purpose.
For each practice, the ITIL maturity model defines criteria for every capability
level from level 2 to level 5. These criteria can be used to assess the practice’s
ability to fulfil its purpose and to contribute to the organization’s service value
system.
Each criterion is mapped to one of the four dimensions of service
management and to the supported capability level. The higher the capability
level, the more comprehensive realization of the practice is expected. For
example, criteria related to the automation of practices are typically defined at
levels 3 or higher because effective automation is only possible if the practice
is well defined and organized.
Figure 7.1 Design of the capability criteria

This approach results in every practice having up to 30 capability criteria based


on the practice PSFs and mapped to the four dimensions of service
management. The number of criteria at each level differs; the four dimensions
are comprehensively covered starting from level 3, so this level typically has
more criteria than others.
Table 7.1 outlines the capability criteria that are defined in the ITIL maturity
model for the problem management practice.
Table 7.1 Problem management capability criteria

PSF Criterion Dimension Capability


level
Identifying and The causes of important Value streams 2
understanding the incidents are identified and and processes
problems and their investigated
impact on services

The errors in products and Value streams 2


services, which might lead and processes
to incidents, are identified
and investigated

Problem identification and Information 3


control address the errors in and technology
technology solutions

Problem identification and Organizations 3


control address the errors and people
within the organization

Problem identification and Partners and 3


control address the errors in suppliers
third-party services and
dependencies

Problem identification and Value streams 3


control address the errors in and processes
value streams, processes,
and other workflows

Problem identification and Value streams 3


control are integrated into and processes
value streams
Problem identification and Information 3
control information is and technology
tracked and managed using
an integrated information
system

The competencies required Organizations 3


to identify and investigate and people
problems are identified and
qualified human resources
are available

Problem identification and Partners and 4


control are integrated suppliers
across the organization’s
supply network

The effectiveness of Value streams 4


problem identification and and processes
control is measured and
reported

The effectiveness of Value streams 5


problem identification and and processes
control is regularly reviewed
and continually improved
Optimizing problem Where reasonably possible, Value streams 2
resolution and errors in products and and processes
mitigation services are fixed

Where errors cannot be Value streams 2


fixed, they are controlled and processes
and mitigated

When error resolutions are Information 3


considered, they address and technology
the technology solutions

When error resolutions are Organizations 3


considered, they address and people
the organization and
competencies

When error resolutions are Partners and 3


considered, they address suppliers
third-party services and
dependencies

When error resolutions are Value streams 3


considered, they address and processes
value streams, processes,
and other workflows

Information about known Information 3


errors is tracked and and technology
managed using an
integrated information
system
Known errors are regularly Value streams 4
reviewed and the relevant and processes
information is updated

The effectiveness of the Value streams 4


error control is measured and processes
and reported

The effectiveness of the Value streams 5


error control is regularly and processes
reviewed and continually
improved

These capability criteria can be used by organizations for self-assessment and


improvement of the practice.

7.2 Capability self-assessment


The self-assessment can be conducted by the service provider’s internal audit
team, if the service provider has one, or by the respective team of the parent
organization. If there is no specialized team in the organization, the
assessment can be done by a team of practice owners and managers
responsible for other management practices of the service provider, or a
mixed team of the service provider’s executive leaders and managers.
To perform a quick self-assessment using the capability criteria, the following
rules should be followed.
1. Start with the level 2 criteria. Based on the knowledge of your
organization, answer the question, ‘Is this a valid description of our
organization in MOST cases?’
2. If the answer to the question above is ‘yes’, make a list of at least
three types of material evidence that could prove the answer. These
can be records, documents, interviews with business stakeholders,
or service provider’s employees.
3. If the answer is ‘yes’ to all criteria of level 2, this level is considered
achieved. Proceed to the criteria of level 3.
4. If not all criteria of level 2 are met, the practice is considered to be at
level 1. Focus on the criteria that are not met; what is missing in the
organization? Why? How can it affect the quality of the IT products
and services? What can be done to meet the criteria that are
currently missed?
5. The same approach is applied at every next level; the practice is
considered to be at the level at which all criteria are met. It is
important to focus on the missing capabilities and improvement
opportunities, rather than on a formal achievement of a high
capability level.

7.3 Problem management capability development


Management practices should support the achievement of the organization’s
objectives and enable creation of value for the stakeholders. Depending on
the service provider’s strategy, positioning, and business and operating
models, some practices may be more important and therefore require a
higher level of capability, however achieving the highest capability level should
never be the aim for all practices. There is no organization that requires all
management practices to be at capability level 5. A higher capability level
provides higher assurance of the fulfilment of a practice’s purpose, but this
also comes with an increase in costs (for instance for management,
automation, or training). To achieve optimal performance with a sufficient level
of assurance, organizations should define a target capability level for each
management practice.
Figure 7.2 and Table 7.2 show the capability development model, which can be
applied to every management practice. The structure of this publication is
aligned with the development steps.
Figure 7.2 The capability development steps and levels

Table 7.2 The problem management capability development steps

Capability Define, agree, Comment for problem Chapter (for


level and implement management recommendations)

2 Purpose and Key sources of problem 2.1


objectives identification
2.3
Scope

Process and 2.2, 3.1


activities Workflows; problem
identification; investigation 4
Roles and and resolutions; roles and
responsibilities responsibilities 5

Tools and
procedures
3 Dependencies Automation and 5
and integration information exchange, use
of an integrated service
management system

Suppliers and other parties 6


involved in problem
management

4 Measurement Metrics 2.5


and reporting

5 Continual Regular review of the 2.4, 2.5, 7


improvement practice and the problem
management capability
development

8. Recommendations for practice success


8.1 Problem management is a practice, not just a set of activities
Successful problem management requires a combination of people with
relevant skills, effective automation, and effective cooperation between
internal teams and with third parties.
A common misconception is that problem management is an administrative
process, whereas in fact it needs collaboration across the teams, enabled by
management involvement and support.
Problem management is a business-led practice, which supports quality,
efficiency, risk reduction and creation of business value.

8.2 Getting started


Start where you are and progress iteratively with feedback. There is no need to
cover all types of problems in all products from day one. Use the capability
criteria (see section 7.2) as guidance. Start identifying problems in the most
critical products and services and then develop the capabilities and expand
the scope to increase business value from the practice.
It is important consider the who the right person or people are to drive the practice
development:
· Do they understand the value and definition of problem
management?
· Do they have the right skills for the relevant phases of the practice?

· What do they need from the leadership to support the practice and
ensure that it is understood across the organization?

8.3 Making problem


management work
Once the first problems have been identified, it is useful to look at data on the
backlog, identify links with incidents and changes, look at other relevant
activities happening in the organization, and seek out feedback from users on
what is really affecting them. Context is important to fully understand the
problems.
Business value is the key driver for problem investigation and resolution.
Management support is often needed to approve investment or resource
allocation. This requires a clear definition of the risks, opportunities, costs and
benefits, not simply the technical issues. Problem management should be
able to present this information clearly and simply to the decision makers.
Visibility and transparency of problems and their impact is also
recommended. The more people across an organization can see what the top
problems are, the more chances that their previous experience may be useful
for the investigation and resolution.
A simple way to test whether problem management is working is to check if
the prioritized list of problems is known and agreed upon by key stakeholders,
particularly the leadership team.

8.4 Demonstrating value to the organization


It is not uncommon for organizations to have defined and implemented a
problem management workflow, and yet still not be able to demonstrate its
business value. This highlights the fact that success in problem management
requires focus on people, collaboration, and business focus, as much as on
process and technology.
Although there are useful metrics that can be used to evaluate the
performance of the practice (see section 2.5), its value should be articulated in
business terms:
· Showing the decrease in costs and losses from the reduction of
incidents and service interruptions.
· Identifying the reduced risk levels associated with the prevention of
incidents and resolution of problems.
· Demonstrating improvements in customer/user satisfaction ratings
due to reduced interruptions and improved service quality.
· Highlighting the quality improvements in service delivery and
performance achieved through problem resolution.
· Showing the return on investments in upgrades, legacy removals,
replacements, resource changes, and use of external services.
It is important to show the value of problem management to the organization
in order to maintain support and investment in the practice.
Most of the content of the practice guides should be taken as a suggestion of
areas that an organization might consider when establishing and nurturing
their own practices. The practice guides are catalogues of topics that
organizations might think about, not a list of answers. When using the
content of the practice guides, organizations should always follow the ITIL
guiding principles:
· focus on value
· start where you are
· progress iteratively with feedback
· collaborate and promote visibility
· think and work holistically
· keep it simple and practical
· optimize and automate.
Table 8.1 outlines recommendations for the success of the problem
management practice, linked to the relevant guiding principles.
Table 8.1 Recommendations for the success of problem management

Recommendation Comments ITIL guiding


principles
Start logging problems, This starts to build up data Start where
now you are
This also gets people to think about long
term problems and remove barriers to Keep it
getting them raised and fixed. They might, simple and
for example, be concerned that they will practical
have to take on the problem, and so not
log it.

Think carefully about The people doing problem management Focus on


getting the right people roles and tasks need to have the right skills value
in problem roles and attributes for the relevant phases
Think and
This includes work
holistically
Data analytics, reporting, trends analytics,
information presentation Keep it
simple and
Risk analysis, business case analysis practical

Project management, leadership,


influencing and communications skills

Problem management Some administrative skills are required to Think and


is not an administrative manage work and promote issues, but this work
function is not the main function or value of holistically
problem management
Keep it
Problem management wont simply work simple and
by itself, particularly when starting the practical
practice; this needs drive and focus
The problem manager Problems need to be owned and acted Collaborate
or problem team won’t upon by the whole and promote
fix all the problems business/team/department. visibility

Problem management facilitates Keep it


(business) fixes rather than delivering all simple and
(technical) fixes practical

Those raising problems are not necessarily


the same as those who will fix them

Prioritize problems in Risk, opportunity, cost and quality should Focus on


order of value to the drive prioritization not technical or team value
organization focus.
Keep it
Problems should be highlighted in simple and
business terms for approval/decision on practical
action

Publish a list of top Transparency of current high priority Focus on


business problems problems ensures focus and shared goals value

Everyone should know the list of top Collaborate


problems and promote
visibility
Open visibility also helps to use the
cumulative experience across a team to Keep it
speed up ideas and actions for resolution simple and
practical
Problem Management Problems will often be raised through the Focus on
needs a swarming tiered support model, but this is not always value
approach effective when collaboration and
management support is needed Collaborate
and promote
All technical teams should be expected to visibility
participate in collaborative problem-
solving activities and swarm teams

SLAs don’t apply to Problems are not predictable and don’t Focus on
problems, but problem always follow a linear process path value
management improves
service quality Problem management should be Progress
measured and appreciated as a iteratively
contributor to business value, for reporting with
feedback
Problem management can also be used to
drive improvement targets for incident Start where
reduction you are

Keep it
simple and
practical

Get user/customer It is essential to get end user feedback on Progress


input on their problems their problems (which also might not be iteratively
defined or logged) with
feedback
It is also beneficial to get an external view
on what really matters to the business Focus on
value

Keep it
simple and
practical
Use AI and automation AI can interrogate data to identify trends Optimize
tools where possible and
For example, AI and RPA can be used to automate
remove repeat and underlying issues

9. Glossary
four dimensions of service management
The four perspectives that are critical to the effective and efficient facilitation
of value for customers and other stakeholders in the form of products and
services.
information and technology
One of the four dimensions of service management. It includes the
information and knowledge used to deliver services, and the information and
technologies used to manage all aspects of the service value system.
ITIL continual improvement model
A model which provides organizations with a structured approach to
implementing improvements.
ITIL guiding principles
Recommendations that can guide an organization in all circumstances,
regardless of changes in its goals, strategies, type of work, or management
structure.
ITIL maturity model
A tool that organizations can use to objectively and comprehensively assess
their service management capabilities and the maturity of their service value
system.
ITIL service value chain
An operating model for service providers that covers all the key activities
required to effectively manage products and services.
known error
A problem that has been analysed but has not been resolved.
metric
A measurement or calculation that is monitored or reported for management
and improvement.
organization
A person or a group of people that has its own functions with responsibilities,
authorities, and relationships to achieve its objectives.
organizations and people
One of the four dimensions of service management. It ensures that the way an
organization is structured and managed, as well as its roles, responsibilities,
and systems of authority and communication, is well defined and supports its
overall strategy and operating model.
output
A tangible or intangible deliverable of an activity.
partners and suppliers
One of the four dimensions of service management. It encompasses the
relationships an organization has with other organizations that are involved in
the design, development, deployment, delivery, support, and/or continual
improvement of services.
practice
A set of organizational resources designed for performing work or
accomplishing an objective. These resources are grouped into the four
dimensions of service management.
practice success factor
A complex functional component of a practice that is required for the practice
to fulfil its purpose.
prioritization
The action of selecting which tasks to work on first when it is impossible to
assign resources to all tasks in the backlog.
problem
A cause, or potential cause, of one or more incidents.
problem model
A repeatable approach to the management of a particular type of problem.
process
A set of interrelated or interacting activities that transform inputs into outputs.
A process takes one or more defined inputs and turns them into defined
outputs. Processes define the sequence of actions and their dependencies.
service provider
A role performed by an organization in a service relationship to provide
services to consumers.
service provision
Activities performed by an organization to provide services and/or supply
goods. Service provision includes:
· management of the provider’s resources, configured to deliver the
service
· ensuring access to these resources for users
· fulfilment of the agreed service actions
· service level management and continual improvement.
service relationship
A cooperation between a service provider and service consumer. Service
relationships include service provision, service consumption, and service
relationship management. Relationships can be basic, cooperative or
collaborative (also known as a partnership).
service value system
A model representing how all the components and activities of an
organization work together to facilitate value creation.
stakeholder
A person or organization that has an interest or involvement in an
organization, product, service, practice, or other entity.
supplier
A stakeholder responsible for providing services that are used by an
organization.
swarming
A technique for solving various complex tasks. In swarming, multiple people
with different areas of expertise work together on a task until it becomes clear
which competencies are the most relevant and needed.
task priority
The importance of a task relative to other tasks. Tasks with a higher priority
should be worked on first. Priority is defined in the context of all tasks in a
backlog.
technical debt
The total rework backlog accumulated by choosing workarounds instead of
systemic solutions that would take longer.
user
A person who uses services.
value
The perceived benefits, usefulness, and
importance of something.
value stream
A series of steps an organization undertakes to create and deliver products
and services to consumers.
value streams and processes
One of the four dimensions of service management. It defines the activities,
workflows, controls, and procedures needed to achieve the agreed objectives.
workaround
A solution that reduces or eliminates the impact of an incident or problem for
which a full resolution is not yet available. Some workarounds reduce the
likelihood of incidents.

10. Acknowledgements
PeopleCert is grateful to everyone who has contributed to the development of
this Official Practice Guide. These Official Practice Guides incorporate an
unprecedented level of enthusiasm and feedback from across the ITIL
community. In particular, PeopleCert would like to thank the following people
Authors

Barry Corless, Roman Zhuravlev, Andrew Vermes

Reviewers
James Ainsworth, Akshay Anand, Sofi Fahlberg, Michael G. Hall, Steve Harrop, Piia
Karvonen, Anton Lykov, Paula Määttänen, Caspar Miller, Christian F. Nissen, Mark
O’Loughlin, Tatiana Orlova, Elina Pirjanti, Stuart Rance

2023 Revision

David Cannon, Antonina Douannes, Peter Farenden, Adam Griffith, Roman Zhuravlev,
Kaimar Karu, Barclay Rae, Stuart Rance, Nicola Reeves

>

Professionals c

Partners c

About us c

Support c

FOLLOW US ON c

Common questions

Powered by AI

Reactive problem identification uses information about past and ongoing incidents to investigate their causes, often prompted by an incident investigation that couldn't identify the incident's nature. It is urgent and focuses on grouping incidents with common causes and business impact . In contrast, proactive problem identification involves regular reviews of incident records for a system, service, or product. This approach seeks to prevent potential issues by identifying causes proactively, especially when there are recurring patterns of incidents or major incidents .

Localization is enhanced by identifying and analyzing configuration items potentially causing incidents. It involves further diagnostics to pinpoint errors within those items. Subsequent activities can include reassignment to teams with the required expertise, thus allowing deeper investigation and addressing errors in doubtful CIs .

Organizations demonstrate the value of problem management by showing reductions in costs and disruptions due to fewer incidents, highlighting improvements in service delivery and customer satisfaction, and presenting the return on investment for changes like upgrades or resource adjustments. Effective presentation of risks, opportunities, and benefits ensures management support and clarifies alignment with business goals .

Key activities in the problem control process include problem investigation to diagnose the root causes of incidents and errors in configuration items (CIs), communication of known errors to relevant teams, and monitoring the status of unresolved known errors to ensure their impact on service is minimized . Problem control may also involve developing solutions and initiating resolutions for known errors and recording actions and results in the problem records .

The impact of reactive problem identification is assessed through statistical analysis, impact analysis, and trend analysis of past incidents. This process evaluates the contribution of a group of incidents to the overall business impact, helping prioritize problems that warrant investigation and potential resolution .

Triggers for registering a problem record include a high number of similar incidents, major incidents, incidents resolved beyond the target resolution time, and availability falling below the target level. These indicators signal the need for investigation into underlying causes and are reviewed by specialist teams responsible for relevant systems, services, or products .

Problem management aligns with creating business value by reducing disruptions, improving service quality, and decreasing risks through effective problem resolution. It also involves clear presentation of technical issues in business terms to secure support for resource allocation, ensuring a direct connection between problem management activities and business objectives .

Successful development of problem management capabilities requires defining clear objectives, integrating the practice within the organization’s service management system, and focusing on critical products first. Ensuring collaboration across teams, enhancing automation where appropriate, and continually assessing the metrics for practice performance are crucial strategies .

Problem initiators play a role in the reactive problem identification process by potentially receiving notifications about the registration of a problem they identified. They contribute to the initial problem categorization and assignment by providing key inputs, such as descriptions, estimated impacts, and associated affected services .

The ITIL maturity model assesses the capability levels of problem management practices from level 2 to level 5. Each level is defined by specific criteria tied to the four dimensions of service management: information and technology, partners and suppliers, value streams and processes, and organizations and people. Criteria related to automation and integration typically indicate higher levels, reflecting more structured and comprehensive practice implementations .

You might also like