0% found this document useful (0 votes)
18 views24 pages

Disaster Recovery Planning Overview

The document outlines the importance of evaluating IT continuity and resilience, focusing on disaster recovery planning (DRP) and compliance requirements. It details recovery strategies, objectives, and alternatives, emphasizing the need for regular testing and documentation. Additionally, it highlights the significance of application and data storage resiliency, as well as telecommunications protection in maintaining critical business processes during disruptions.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views24 pages

Disaster Recovery Planning Overview

The document outlines the importance of evaluating IT continuity and resilience, focusing on disaster recovery planning (DRP) and compliance requirements. It details recovery strategies, objectives, and alternatives, emphasizing the need for regular testing and documentation. Additionally, it highlights the significance of application and data storage resiliency, as well as telecommunications protection in maintaining critical business processes during disruptions.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Task 4.

10

Evaluate IT continuity and resilience


(backups/restores, disaster recovery plan
[DRP]) to determine whether they are
controlled effectively and continue to support
the organization’s objectives.
Disaster Recovery Planning
§ Planning for disasters is an important part of the risk
management and BCP processes.
§ The purpose of this continuous planning process is to
ensure that cost-effective controls are in place to prevent
possible IT disruptions and to recover the IT capacity of
the organization in the event of a disruption.
DRP Compliance Requirements
§ DRP may be subject to compliance requirements depending
on:
o Geographic location
o Nature of the business
o The legal and regulatory framework
§ Most compliance requirements focus on ensuring continuity of
service with human safety as the most essential objective.
§ Organizations may engage third parties to perform
DRP-related activities on their behalf; these third parties are
also subject to compliance.
Disaster Recovery Testing
§ The IS auditor should ensure that all plans are regularly
tested and be aware of the testing schedule and tests to be
conducted for all critical functions.
§ Test documentation should be reviewed by the IS auditor to
confirm that tests are fully documented with pre-test, test and
post-test reports.
o It is also important that information security is validated to
ensure that it is not compromised during testing.
RPO and RTO Defined

Recovery point objective Recovery time objective


(RPO) (RTO)

• Determined based on the • The amount of time allowed for


acceptable data loss in case of a the recovery of a business
disruption of operations. It function or resource after a
indicates the earliest point in time disaster occurs.
that is acceptable to recover the
data.
• The RPO effectively quantifies
the permissible amount of data
loss in case of interruption.
RPO and RTO Responses
§ Both RPO and RTO are based on time parameters. The nearer the
time requirements are to the center, the more costly the recovery
strategy. Note the strategies employed at each time mark in the
graphic below.

Recovery Point Objective Recovery Time Objective


n
tio
up
isr
D
4-24 hrs 1-4 hrs 0-1 hr 0-1 hr 1-4 hrs 4-24 hrs
• Tape backups • Disk-based • Mirroring • Active-active • Active-passive • Cold standby
• Log shipping backups • Real-time clustering clustering
• Snapshots replication • Hot standby
• Delayed
replication
• Log shipping
Additional Parameters
§ The following parameters are also important in defining recovery
strategies:
o Interruption window—The maximum period of time an
organization can wait from point of failure to critical services
restoration, after which progressive losses from the interruption
cannot be afforded.
o Service delivery objective (SDO)—Directly related to business
needs, this defines the level of services that must be reached
during the alternate processing period.
o Maximum tolerable outages—The amount of time the
organization can support processing in the alternate mode, after
which new problems can arise from lower than usual SDO, and
the accumulation of information pending update becomes
unmanageable.
Recovery Strategies
§ Documented recovery procedures ensure a return to
normal system operations in the event of an interruption.
§ These are based on recovery strategies, which should
be:
o Recommended to and selected by senior
management
o Used to further develop the business continuity plan
(BCP)
Recovery Strategies (cont’d)
§ The selection of a recovery strategy depends on the criticality
of the business process and its associated applications, cost,
security and time to recover.
§ In general, each IT platform running an application that
supports a critical business function will need a recovery
strategy.
§ Appropriate strategies are those in which the cost of recovery
within a specific time frame is balanced by the impact and
likelihood of an occurrence.
§ The cost of recovery includes both the fixed costs of providing
redundant or alternate resources and the variable costs of
putting these into use should a disruption occur.
Recovery Alternatives

Hot sites

• A facility with all of the IT and communications equipment required


to support critical applications, along with office accommodations
for personnel.
Recovery Alternatives (cont’d)

Warm sites
• A complete infrastructure, partially configured for IT, usually with
network connections and essential peripheral equipment. Current
versions of programs and data would likely need to be installed
before operations could resume at the recovery site.

Cold sites

• A facility with the space and basic infrastructure to support the


resumption of operation but lacking any IT or communications
equipment, programs, data or office support.
Recovery Alternatives (cont’d)

Mirrored sites
• A fully redundant site with real-time data replication from the
production site.

Mobile sites

• Modular processing facilities mounted on transportable vehicles,


ready to be delivered and set up on an as-needed basis.
Recovery Alternatives (cont’d)

Reciprocal arrangements

• Agreements between separate, but similar, companies to


temporarily share their IT facilities in the event that a partner to the
agreement loses processing capability.

Reciprocal arrangements with other organizations

• Agreements between two or more organizations with unique


equipment or applications. Participants promise to assist each
other during an emergency.
Application Resiliency
§ The ability to protect an application against a disaster
depends on providing a way to restore it as quickly as
possible.
§ A cluster is a type of software installed on every server in
which an application runs. It includes management
software that permits control of and tuning of the cluster
behavior.
Application Resiliency (cont’d)
§ Clustering protects against single points of failure in
which the loss of a resource would result in the loss of
service or production.
§ There are two major types of application clusters, active-
passive and active-active.
Data Storage Resiliency
§ The data protection method known as RAID, or
Redundant Array of Independent (or Inexpensive) Disks,
is the most common and basic method used to protect
data against loss at a single point of failure.
§ Such storage arrays provide data replication features,
ensuring that the data saved to a disk on one site
appears on the other site.
Data Storage Resiliency (cont’d)
§ Data replication may be:
o Synchronous—Local disk write is confirmed upon data
replication at other site.
o Asynchronous—Data are replicated on a scheduled
basis.
o Adaptive—Switching between synchronous and
asynchronous depending on network load.
Telecommunications Resiliency
§ The DRP should also contain the organization’s
telecommunication networks.
§ These are susceptible to the same interruptions as data
centers and several other issues, for example:
o Central switching office disasters
o Cable cuts
o Security breaches
§ To provide for the maintenance of critical business processes,
telecommunications capabilities must be identified for various
thresholds of outage.
Network Protection

Alternative Diverse
Redundancy
routing routing

Long-haul Last-mile
Voice
network circuit
recovery
diversity protection
Offsite Library Controls

Secure physical Location of the library


access to library Ensuring that the away from the data
Encryption of backup physical construction
contents, accessible media, especially center and disasters
only to authorized can withstand heat, that may strike both
during transit fire and water
persons together

Maintenance of an
inventory of all Maintenance and
Maintenance of library protection of a catalog
storage media and records for specified
files for specified of information
retention periods regarding data files
retention periods
Discussion Question
During an IS audit of the disaster recovery plan (DRP) of a
global enterprise, the IS auditor observes that some remote
offices have very limited local IT resources. Which of the
following observations would be the MOST critical for the IS
auditor?
A. A test has not been made to ensure that local resources
could maintain security and service standards when
recovering from a disaster or incident.
B. The corporate business continuity plan (BCP) does not
accurately document the systems that exist at remote
offices.
C. Corporate security measures have not been incorporated
into the test plan.
D. A test has not been made to ensure that tape backups
from the remote offices are usable.
Discussion Question
Which of the following is the BEST indicator of the
effectiveness of backup and restore procedures while
restoring data after a disaster?
A. Members of the recovery team were available.
B. Recovery time objectives (RTOs) were met.
C. Inventory of backup tapes was properly maintained.
D. Backup tapes were completely restored at an
alternate site.
Domain 4 Summary
§ Evaluate IT service management framework and
practices.
§ Evaluate IT operations (e.g., job scheduling,
configuration management, capacity and performance
management).
§ Evaluate IT maintenance (patches, upgrades).
§ Evaluate database management practices.
Domain 4 Summary (cont’d)
§ Evaluate data quality and life cycle management.
§ Evaluate problem and incident management practices.
§ Evaluate change and release management practices.
§ Evaluate end-user computing.
§ Evaluate IT continuity and resilience (backups/restores,
disaster recovery plan [DRP]).

You might also like