DB2 Database Performance Guide 8.2
DB2 Database Performance Guide 8.2
SC09-4821-01
®
IBM DB2 Universal Database
™
SC09-4821-01
Before using this information and the product it supports, be sure to read the general information under Notices.
This document contains proprietary information of IBM. It is provided under a license agreement and is protected
by copyright law. The information contained in this publication does not include any product warranties, and any
statements provided in this manual should not be interpreted as such.
You can order IBM publications online or through your local IBM representative.
v To order publications online, go to the IBM Publications Center at [Link]/shop/publications/order
v To find your local IBM representative, go to the IBM Directory of Worldwide Contacts at
[Link]/planetwide
To order DB2 publications from DB2 Marketing and Sales in the United States or Canada, call 1-800-IBM-4YOU
(426-4968).
When you send information to IBM, you grant IBM a nonexclusive right to use or distribute the information in any
way it believes appropriate without incurring any obligation to you.
© Copyright International Business Machines Corporation 1993 - 2004. All rights reserved.
US Government Users Restricted Rights – Use, duplication or disclosure restricted by GSA ADP Schedule Contract
with IBM Corp.
Contents
About this book . . . . . . . . . . . ix Correcting lock escalation problems . . . . . 55
Who should use this book . . . . . . . . . . x Evaluate uncommitted data via lock deferral . . 56
How this book is structured . . . . . . . . . x Lock type compatibility . . . . . . . . . 59
A brief overview of the other Administration Guide Lock modes and access paths for standard tables 60
volumes . . . . . . . . . . . . . . . xi Lock modes for table and RID index scans of
Administration Guide: Planning . . . . . . xi MDC tables . . . . . . . . . . . . . 62
Administration Guide: Implementation . . . . xii Locking for block index scans for MDC tables . . 65
Factors that affect locking. . . . . . . . . . 68
Factors that affect locking. . . . . . . . . . 68
Part 1. Introduction to performance 1 Locks and types of application processing . . . 68
Locks and data-access methods . . . . . . . 69
Chapter 1. Introduction to performance 3 Index types and next-key locking . . . . . . 70
Elements of performance . . . . . . . . . . 3 Optimization factors . . . . . . . . . . . 71
Performance tuning guidelines . . . . . . . . 3 Optimization class guidelines . . . . . . . 72
The performance tuning process . . . . . . . . 5 Optimization classes . . . . . . . . . . 73
Developing a performance improvement process . 5 Setting the optimization class . . . . . . . 76
Performance information that users can provide . 6 Tuning applications . . . . . . . . . . . . 77
Performance tuning limits . . . . . . . . . 6 Guidelines for restricting select statements . . . 77
Quick-start tips for performance tuning . . . . . 7 Specifying row blocking to reduce overhead . . 80
Query tuning guidelines . . . . . . . . . 81
Chapter 2. Architecture and processes . 9 Data sampling in SQL queries . . . . . . . 82
DB2 architecture and process overview . . . . . 9 Efficient SELECT statements . . . . . . . . 83
Deadlocks between applications . . . . . . . 11 Compound SQL guidelines . . . . . . . . 85
Disk storage overview . . . . . . . . . . . 12 Character-conversion guidelines . . . . . . 86
Disk-storage performance factors . . . . . . 12 Guidelines for stored procedures . . . . . . 87
Database directories and files . . . . . . . 12 Parallel processing for applications . . . . . 88
Table space overview . . . . . . . . . . . 14 | Improving performance by binding with REOPT 89
SMS table spaces . . . . . . . . . . . 14
DMS table spaces . . . . . . . . . . . 15 Chapter 4. Environmental
Illustration of the DMS table-space address map 17 considerations . . . . . . . . . . . 91
Tables and indexes . . . . . . . . . . . . 18 Database partition group impact on query
Table and index management for standard tables 18 optimization . . . . . . . . . . . . . . 91
Table and index management for MDC tables . . 21 Table space impact on query optimization . . . . 91
Index structure . . . . . . . . . . . . 23 Server options affecting federated databases . . . 94
Processes . . . . . . . . . . . . . . . 25
Log processing . . . . . . . . . . . . 25 Chapter 5. System catalog statistics . . 95
Insert processing . . . . . . . . . . . 26
Catalog statistics . . . . . . . . . . . . . 95
Update processing . . . . . . . . . . . 27
Collecting and analyzing catalog statistics . . . . 96
Client-server processing model . . . . . . . 28
Guidelines for collecting and updating statistics 97
Memory management . . . . . . . . . . 32
Collecting catalog statistics . . . . . . . . 98
Collecting distribution statistics for specific
Part 2. Tuning application columns . . . . . . . . . . . . . . 99
performance . . . . . . . . . . . 37 Collecting index statistics . . . . . . . . 100
| Collecting statistics on a sample of the table
| data . . . . . . . . . . . . . . . 101
Chapter 3. Application considerations 39 | Collecting statistics using a statistics profile . . 102
Concurrency control and isolation levels . . . . . 39 | Automatic statistics collection . . . . . . . 104
Concurrency issues . . . . . . . . . . . 39 | Using automatic statistics collection . . . . . 105
Performance impact of isolation levels . . . . 40 Statistics collected . . . . . . . . . . . . 106
Specifying the isolation level . . . . . . . 43 Catalog statistics tables . . . . . . . . . 106
Concurrency control and locking . . . . . . . 46 Statistical information that is collected . . . . 111
Locks and concurrency control . . . . . . . 46 Distribution statistics . . . . . . . . . . 112
Lock attributes . . . . . . . . . . . . 47 Optimizer use of distribution statistics . . . . 115
Locks and performance . . . . . . . . . 49 Extended examples of distribution-statistics use 116
Guidelines for locking . . . . . . . . . . 53 Detailed index statistics . . . . . . . . . 120
Contents v
smtp_server - SMTP server . . . . . . . . 484 Example two: single-partition plan with
toolscat_db - Tools catalog database . . . . . 485 intra-partition parallelism . . . . . . . . 578
toolscat_inst - Tools catalog database instance 485 Example three: multipartition plan with
toolscat_schema - Tools catalog database schema 486 inter-partition parallelism . . . . . . . . 579
Example four: multipartition plan with
inter-partition and intra-partition parallelism . . 582
Part 4. Appendixes . . . . . . . . 487 Example five: federated database plan . . . . 584
Contents vii
viii Administration Guide: Performance
About this book
The Administration Guide in its three volumes provides information necessary to
use and administer the DB2 relational database management system (RDBMS)
products, and includes:
v Information about database design (found in Administration Guide: Planning)
v Information about implementing and managing databases (found in
Administration Guide: Implementation)
v Information about configuring and tuning your database environment to
improve performance (found in Administration Guide: Performance)
Many of the tasks described in this book can be performed using different
interfaces:
v The Command Line Processor, which allows you to access and manipulate
databases from a graphical interface. From this interface, you can also execute
SQL statements and DB2 utility functions. Most examples in this book illustrate
the use of this interface. For more information about using the command line
processor, see the Command Reference.
v The application programming interface, which allows you to execute DB2
utility functions within an application program. For more information about
using the application programming interface, see the Administrative API
Reference.
| v The Control Center, which allows you to use a graphical user interface to
| perform administrative tasks such as configuring the system, managing
| directories, backing up and recovering the system, scheduling jobs, and
| managing media. The Control Center also contains Replication Administration,
| which allows you set up the replication of data between systems. Further, the
| Control Center allows you to execute DB2 utility functions through a graphical
| user interface. There are different methods to invoke the Control Center
| depending on your platform. For example, use the db2cc command on a
| command line, select the Control Center icon from the DB2 folder, or use the
| Start menu on Windows platforms. For introductory help, select Getting started
| from the Help pull-down of the Control Center window. The Visual Explain
| tool is invoked from the Control Center.
| The Control Center is available in three views:
| – Basic. This view shows the core DB2 UDB functions on essential objects such
| as databases, tables, and stored procedures.
| – Advanced. This view has all of the objects and actions available. Use this
| view if you are working in an enterprise environment and you want to
| connect to DB2 for z/OS or IMS.
| – Custom. This view gives you the ability to tailor the object tree and the object
| actions.
There are other tools that you can use to perform administration tasks. They
include:
| v The Command Editor which replaces the Command Center and is used to
| generate, edit, run, and manipulate SQL statements; IMS and DB2 commands;
| work with the resulting output; and to view a graphical representation of the
| access plan for explained SQL statements.
Introduction to Performance
v Chapter 1, “Introduction to performance,” introduces concepts and
considerations for managing and improving DB2 UDB performance.
v Chapter 2, “Architecture and processes,” introduces underlying DB2 Universal
Database architecture and processes.
Appendixes
v Appendix A, “DB2 Registry and Environment Variables,” presents profile
registry values and environment variables.
v Appendix B, “Explain tables,” The explain table section provides information
about the tables used by the DB2 Explain facility and how to create those tables.
v Appendix C, “SQL explain tools,” provides information on using the DB2
explain tools: db2expln and dynexpln.
v Appendix D, “db2exfmt - Explain Table Format,” formats the contents of the
DB2 explain tables.
Database Concepts
v ″Basic relational database concepts″ presents an overview of database objects,
including recovery objects, storage objects, and system objects.
v ″Parallel database systems″ provides an introduction to the types of parallelism
available with DB2.
v ″About data warehousing″ provides an overview of data warehousing and data
warehousing tasks.
Database Design
v ″Logical database design″ discusses the concepts and guidelines for logical
database design.
v ″Physical database design″ discusses the guidelines for physical database design,
including considerations related to data storage.
v ″Designing distributed databases″ discusses how you can access multiple
databases in a single transaction.
v ″Designing for Transaction Managers″ discusses how you can use your databases
in a distributed transaction processing environment.
Appendixes
v ″Incompatibilities between releases″ presents the incompatibilities introduced by
Version 7 and Version 8, as well as future incompatibilities that you should be
aware of.
v ″National language support (NLS)″ describes DB2 National Language Support,
including information about territories, languages, and code pages.
v ″Enabling large page support in a 64-bit environment (AIX)″ discusses the
support for a 16 MB page size and how to enable this support.
Database Security
v ″Controlling Database Access″ describes how you can control access to your
database’s resources.
v ″Auditing DB2 Activities″ describes how you can detect and monitor unwanted
or unanticipated access to data.
Appendixes
| v ″Conforming to the naming rules″ presents the rules to follow when naming
| databases and objects.
| v ″Using automatic client rerouting″ discusses the automatic rerouting of client
| applications and how to enable this support.
| v ″Using lightweight directory access protocol (LDAP) directory services″ provides
| information about how you can use LDAP Directory Services.
| v ″Issuing commands to multiple database partitions″ discusses the use of the
| db2_all and rah shell scripts to send commands to all partitions in a partitioned
| database environment.
| v ″Windows Management Instrumentation (WMI) support″ describes how DB2
| supports this management infrastructure standard to integrate various hardware
| and software management systems. Also discussed is how DB2 integrates with
| WMI.
| v ″Using Windows NT security″ describes how DB2 Universal Database works
| with Windows NT security.
v ″Using the Windows Performance Monitor″ provides information about
registering DB2 with the Windows NT Performance Monitor, and using the
performance information.
| v ″Using Windows database partition servers″ provides information about the
| utilities available to work with database partition servers on Windows NT or
| Windows 2000.
| v ″Configuring multiple logical nodes″ describes how to configure multiple logical
| nodes in a partitioned database environment.
All of the information on the DB2 utilities for moving data, and the
comparable topics from the Command Reference and the Administrative API
Reference, have been consolidated into the Data Movement Utilities Guide and
Reference.
The Data Movement Utilities Guide and Reference is your primary, single
source of information for these topics.
To find out more about replication of data, see IBM DB2 Information
Integrator SQL Replication Guide and Reference.
All of the information on the methods and tools for backing up and
recovering data, and the comparable topics from the Command Reference and
the Administrative API Reference, have been consolidated into the Data
Recovery and High Availability Guide and Reference.
The Data Recovery and High Availability Guide and Reference is your primary,
single source of information for these topics.
Elements of performance
Performance is the way a computer system behaves given a particular work load.
Performance is measured in terms of system response time, throughput, and
availability. Performance is also affected by:
v The resources available in your system
v How well those resources are used and shared.
In general, you tune your system to improve its cost-benefit ratio. Specific goals
could include:
v Processing a larger, or more demanding, work load without increasing
processing costs
For example, to increase the work load without buying new hardware or using
more processor time
v Obtaining faster system response times, or higher throughput, without
increasing processing costs
v Reducing processing costs without degrading service to your users
Other benefits, such as greater user satisfaction because of quicker response time,
are intangible. All of these benefits should be considered.
Related concepts:
v “Performance tuning guidelines” on page 3
v “Quick-start tips for performance tuning” on page 7
Related tasks:
v “Developing a performance improvement process” on page 5
Do not tune just for the sake of tuning: Tune to relieve identified constraints. If
you tune resources that are not the primary cause of performance problems, this
has little or no effect on response time until you have relieved the major
constraints, and it can actually make subsequent tuning work more difficult. If
there is any significant improvement potential, it lies in improving the performance
of the resources that are major factors in the response time.
Consider the whole system: You can never tune one parameter or system in
isolation. Before you make any adjustments, consider how it will affect the system
as a whole.
Change one parameter at a time: Do not change more than one performance
tuning parameter at a time. Even if you are sure that all the changes will be
beneficial, you will have no way of evaluating how much each change contributed.
You also cannot effectively judge the trade-off you have made by changing more
than one parameter at a time. Every time you adjust a parameter to improve one
area, you almost always affect at least one other area that you may not have
considered. By changing only one at a time, this allows you to have a benchmark
to evaluate whether the change does what you want.
Measure and reconfigure by levels: For the same reasons that you should only
change one parameter at a time, tune one level of your system at a time. You can
use the following list of levels within a system as a guide:
v Hardware
v Operating System
v Application Server and Requester
v Database Manager
v SQL Statements
v Application Programs
Understand the problem before you upgrade your hardware: Even if it seems that
additional storage or processor power could immediately improve performance,
take the time to understand where your bottlenecks are. You may spend money on
additional disk storage only to find that you do not have the processing power or
the channels to exploit it.
Put fall-back procedures in place before you start tuning: As noted earlier, some
tuning can cause unexpected performance results. If this leads to poorer
performance, it should be reversed and alternative tuning tried. If the former setup
is saved in such a manner that it can be simply recalled, the backing out of the
incorrect information becomes much simpler.
Related concepts:
v “Elements of performance” on page 3
v “Quick-start tips for performance tuning” on page 7
Base your performance monitoring and tuning decisions on your knowledge of the
kinds of applications that use the data and the patterns of data access. Different
kinds of applications have different performance requirements.
Procedure:
Related concepts:
v “Elements of performance” on page 3
v “Performance tuning guidelines” on page 3
v “Quick-start tips for performance tuning” on page 7
v “Performance tuning limits” on page 6
v “Performance information that users can provide” on page 6
Related concepts:
v “Performance tuning guidelines” on page 3
Related tasks:
v “Developing a performance improvement process” on page 5
For example, tuning can often improve performance if the system encounters a
performance bottleneck. If you are close to the performance limits of your system
and the number of users increases by about ten percent, the response time is likely
to increase by much more than ten percent. In this situation, you need to
determine how to counterbalance this degradation in performance by tuning your
system.
However, there is a point beyond which tuning cannot help. At this point, consider
revising your goals and expectations within the limits of your environment. For
significant performance improvements, you might need to add more disk storage,
faster CPU, additional CPUs, more main memory, faster communication links, or a
combination of these.
Related concepts:
v “Management of database server capacity” on page 281
Related tasks:
v “Developing a performance improvement process” on page 5
Note: If you use the ACTIVATE DATABASE command, you must shut down the
database with the DEACTIVATE DATABASE command. The last
application that disconnects from the database does not shut it down.
v Consult the summary tables that list and briefly describe each configuration
parameter available for the database manager and each database.
These summary tables contain a column that indicates whether tuning the
parameter results in high, medium,low, or no performance changes, either for
better or for worse. Use this table to find the parameters that you might tune for
the largest performance improvements.
Related concepts:
v “The database system monitor information” on page 262
Related reference:
v “Configuration parameters summary” on page 323
The following figure shows a general overview of the architecture and processes
for DB2 UDB.
Clients
Client Client
application application
UDB Client Library
UDB server
Logical
agents
Log buffer
Write log Coordinator Coordinator
requests agent agent
Async I/O
Subagents Subagents prefetch
requests
Victim
notifications
Logger Deadlock
detector
Scatter/Gather
I/Os Prefetchers
Log Parallel,
big-block,
read requests
On the server side, activity is controlled by engine dispatchable units (EDUs). In all
figures in this section, EDUs are shown as circles or groups of circles. EDUs are
implemented as threads in a single process on Windows®-based platforms and as
processes on UNIX®. DB2 agents are the most common type of EDUs. These agents
perform most of the SQL processing on behalf of applications. Prefetchers and
page cleaners are other common EDUs.
All agents and subagents are managed using a pooling algorithm that minimizes
the creation and destruction of EDUs.
Buffer pools are areas of database server memory where database pages of user
table data, index data, and catalog data are temporarily moved and can be
modified. Buffer pools are a key determinant of database performance because
data can be accessed much faster from memory than from disk. If more of the data
needed by applications is present in a buffer pool, less time is required to access
the data than to find it on disk.
The configuration of the buffer pools, as well as prefetcher and page cleaner EDUs,
controls how quickly data can be accessed and how readily available it is to
applications.
v Prefetchers retrieve data from disk and move it into the buffer pool before
applications need the data. For example, applications needing to scan through
large volumes of data would have to wait for data to be moved from disk into
the buffer pool if there were no data prefetchers. Agents of the application send
asynchronous read-ahead requests to a common prefetch queue. As prefetchers
become available, they implement those requests by using big-block or
scatter-read input operations to bring the requested pages from disk to the
buffer pool. If you have multiple disks for storage of the database data, the data
can be striped across the disks. Striping data lets the prefetchers use multiple
disks at the same time to retrieve data.
v Page cleaners move data from the buffer pool back out to disk. Page cleaners
are background EDUs that are independent of the application agents. They look
for pages from the buffer pool that are no longer needed and write the pages to
disk. Page cleaners ensure that there is room in the buffer pool for the pages
being retrieved by the prefetchers.
Without the independent prefetchers and the page cleaner EDUs, the application
agents would have to do all of the reading and writing of data between the buffer
pool and disk storage.
Related concepts:
v “Prefetching data into the buffer pool” on page 229
v “Deadlocks between applications” on page 11
v “Database directories and files” on page 12
Related reference:
v “max_coordagents - Maximum number of coordinating agents” on page 379
v “max_connections - Maximum number of client connections” on page 379
Deadlock concept
Table 1
Application A .. Application B
T1: update row 1 of table 1 . T1: update row 2 of table 2
T2: update row 2 of table 2 Row x T2: update row 1 of table 1
T3: deadlock
.. 1 T3: deadlock
.
Row 2
..
.
Table 2
..
.
Row.. 1
.
x Row 2
..
.
Because applications do not voluntarily release locks on data that they need, a
deadlock detector process is required to break deadlocks and allow application
processing to continue. As its name suggests, the deadlock detector monitors the
information about agents waiting on locks. The deadlock detector arbitrarily selects
one of the applications in the deadlock and releases the locks currently held by
that “volunteered” application. By releasing the locks of that application, the data
required by other waiting applications is made available for use. The waiting
applications can then access the data required to complete transactions.
Related concepts:
v “Locks and performance” on page 49
Related concepts:
v “DMS device considerations” on page 255
v “Database directories and files” on page 12
v “SMS table spaces” on page 14
v “DMS table spaces” on page 15
v “Table and index management for standard tables” on page 18
v “Table and index management for MDC tables” on page 21
It is recommended that you explicitly state where you would like the database
created.
The database directory contains the following files that are created as part of the
CREATE DATABASE command.
v The files SQLBP.1 and SQLBP.2 contain buffer pool information. Each file has a
duplicate copy to provide a backup.
v The files SQLSPCS.1 and SQLSPCS.2 contain table space information. Each file
has a duplicate copy to provide a backup.
v The SQLDBCON file contains database configuration information. Do not edit
this file. To change configuration parameters, use either the Control Center or
the command-line statements UPDATE DATABASE CONFIGURATION and
RESET DATABASE CONFIGURATION.
v The [Link] history file and its backup [Link] contain history
information about backups, restores, loading of tables, reorganization of tables,
altering of a table space, and other changes to a database.
The [Link] file contains a history of table space changes at a log-file
level. For each log file, [Link] contains information that helps to
identify which table spaces are affected by the log file. Table space recovery uses
information from this file to determine which log files to process during table
space recovery. You can examine the contents of both history files in a text
editor.
v The log control files, [Link] and [Link], contain information
about the active logs.
Recovery processing uses information from this file to determine how far back in
the logs to begin recovery. The SQLOGDIR subdirectory contains the actual log
files.
Note: You should ensure the log subdirectory is mapped to different disks than
those used for your data. A disk problem could then be restricted to your
data or the logs but not both. This can provide a substantial performance
benefit because the log files and database containers do not compete for
movement of the same disk heads. To change the location of the log
subdirectory, change the newlogpath database configuration parameter.
v The SQLINSLK file helps to ensure that a database is used by only one instance
of the database manager.
At the same time a database is created, a detailed deadlocks event monitor is also
created. The detailed deadlocks event monitor files are stored in the database
directory of the catalog node. When the event monitor reaches its maximum
number of files to output, it will deactivate and a message is written to the
notification log. This prevents the event monitor from consuming too much disk
space. Removing output files that are no longer needed will allow the event
monitor to activate again on the next database activation.
In addition, a file called SQL*.DAT stores information about each table that the
subdirectory or container contains. The asterisk (*) is replaced by a unique set of
digits that identifies each table. For each SQL*.DAT file there might be one or more
of the following files, depending on the table type, the reorganization status of the
table, or whether indexes, LOB, or LONG fields exist for the table:
| v SQL*.BKM (contains block allocation information if it is an MDC table)
v SQL*.LF (contains LONG VARCHAR or LONG VARGRAPHIC data)
v SQL*.LB (contains BLOB, CLOB, or DBCLOB data)
v SQL*.LBA (contains allocation and free space information about SQL*.LB files)
v SQL*.INX (contains index table data)
| v SQL*.IN1 (contains index table data)
v SQL*.DTR (contains temporary data for a reorganization of an SQL*.DAT file)
v SQL*.LFR (contains temporary data for a reorganization of an SQL*.LF file)
v SQL*.RLB (contains temporary data for a reorganization of an SQL*.LB file)
v SQL*.RBA (contains temporary data for a reorganization of an SQL*.LBA file)
Related concepts:
v “Comparison of SMS and DMS table spaces” in the Administration Guide:
Planning
v “DMS device considerations” on page 255
v “SMS table spaces” on page 14
v “DMS table spaces” on page 15
v “Illustration of the DMS table-space address map” on page 17
v “Understanding the recovery history file” in the Data Recovery and High
Availability Guide and Reference
Related reference:
v “CREATE DATABASE Command” in the Command Reference
| In an SMS table space, space for tables is allocated on demand. The amount of
| space that is allocated is dependent on the setting of the multipage_alloc database
| configuration parameter. If this configuration parameter is set to YES, then a full
| extent will be allocated when space is required. Otherwise, space will be allocated
| one page at a time. Prior to version 8.2, the default setting of the configuration
| parameter was NO which caused one page to be allocated at a time. This default
| could be changed with the db2empfa tool. When you run db2empfa, the
| multipage_alloc database configuration parameter is set to YES. In version 8.2, the
| default setting of the configuration parameter is set to YES which means that a full
| extent is allocated at a time by default.
| Multi-page file allocation only affects the data and index portions of a table. This
| means that the .LF, .LB, and .LBA files are not extended one extent at a time.
When all space in a single container in an SMS table space is allocated to tables,
the table space is considered full, even if space remains in other containers. You
can add containers to an SMS table space only on a partition that does not yet
have any containers.
Note: SMS table spaces can take advantage of file-system prefetching and caching .
Related concepts:
v “Table space design” in the Administration Guide: Planning
v “Comparison of SMS and DMS table spaces” in the Administration Guide:
Planning
Related tasks:
v “Adding a container to an SMS table space on a partition” in the Administration
Guide: Implementation
Related reference:
v “multipage_alloc - Multipage file allocation enabled” on page 429
v “db2empfa - Enable Multipage File Allocation Command” in the Command
Reference
DMS table spaces differ from SMS table spaces in that for DMS table spaces, space
is allocated when the table space is created and not allocated when needed.
Also, placement of data can differ on the two types of table spaces. For example,
consider the need for efficient table scans: it is important that the pages in an
extent are physically contiguous. With SMS, the file system of the operating system
Note: Like SMS table spaces, DMS file containers can take advantage of file-system
prefetching and caching. However, DMS table spaces cannot.
Unlike SMS table spaces, the containers that make up a DMS table space do not
need to be close to being equal in their capacity. However, it is recommended that
the containers are equal, or close to being equal, in their capacity. Also, if any
container is full, any available free space from other containers can be used in a
DMS table space.
When working with DMS table spaces, you should consider associating each
container with a different disk. This allows for a larger table space capacity and the
ability to take advantage of parallel I/O operations.
The CREATE TABLESPACE statement creates a new table space within a database,
assigns containers to the table space, and records the table space definition and
attributes in the catalog. When you create a table space, the extent size is defined
as a number of contiguous pages. The extent is the unit of space allocation within
a table space. Only one table or other object, such as an index, can use the pages in
any single extent. All objects created in the table space are allocated extents in a
logical table space address map. Extent allocation is managed through Space Map
Pages (SMP).
The first extent in the logical table space address map is a header for the table
space containing internal control information. The second extent is the first extent
of Space Map Pages (SMP) for the table space. SMP extents are spread at regular
intervals throughout the table space. Each SMP extent is simply a bit map of the
extents from the current SMP extent to the next SMP extent. The bit map is used to
track which of the intermediate extents are in use.
The next extent following the SMP is the object table for the table space. The object
table is an internal table that tracks which user objects exist in the table space and
where their first Extent Map Page (EMP) extent is located. Each object has its own
EMPs which provide a map to each page of the object that is stored in the logical
table space address map.
Related concepts:
v “Table space design” in the Administration Guide: Planning
v “Comparison of SMS and DMS table spaces” in the Administration Guide:
Planning
Related tasks:
v “Adding a container to a DMS table space” in the Administration Guide:
Implementation
Related reference:
v “CREATE TABLESPACE statement” in the SQL Reference, Volume 2
The object table is an internal relational table that maps an object identifier to the
location of the first EMP extent in the table. This EMP extent, directly or indirectly,
maps out all extents in the object. Each EMP contains an array of entries. Each
entry maps an object-relative extent number to a table space-relative page number
where the object extent is located. Direct EMP entries directly map object-relative
addresses to table space-relative addresses. The last EMP page in the first EMP
extent contains indirect entries. Indirect EMP entries map to EMP pages which
then map to object pages. The last 16 entries in the last EMP page in the first EMP
extent contain double-indirect entries.
The extents from the logical table-space address map are striped in round-robin
order across the containers associated with the table space.
1 4021
C RID
K RID
2 4022
3 4023
RID (record ID) = Page 4023, Slot 2
Legend
... ...
... ... reserved for system records
FSCR
user records
Figure 4. Logical table, record, and index structure for standard tables
In standard tables, data is logically organized as a list of data pages. These data
pages are logically grouped together based on the extent size of the table space.
The number of records contained within each data page can vary based on the size
of the data page and the size of the records. A maximum of 255 records can fit on
one page. Most pages contain only user records. However, a small number of
pages include special internal records, that are used by DB2® to manage the table.
For example, in a standard table there is a Free Space Control Record (FSCR) on
every 500th data page. These records map the free space for new records on each
of the following 500 data pages (until the next FSCR). This available free space is
used when inserting records into the table.
Logically, index pages are organized as a B-tree which can efficiently locate records
in the table that have a given key value. The number of entities on an index page
is not fixed but depends on the size of the key. For tables in DMS table spaces,
record identifiers (RIDs) in the index pages use table space-relative page numbers,
not object-relative page numbers. This allows an index scan to directly access the
data pages without requiring an Extent Map page (EMP) for mapping.
Each data page has the same format. A page header begins each data page. After
the page header there is a slot directory. Each entry in the slot directory
corresponds to a different record on the page. The entry itself is the byte-offset into
the data page where the record begins. Entries of minus one (-1) correspond to
deleted records.
Record identifiers (RIDs) are a three-byte page number followed by a one-byte slot
number. Type-2 index records also contain an additional byte called the ridFlag.
The ridFlag stores information about the status of keys in the index, such as
whether this key has been marked deleted. Once the index is used to identify a
RID, the RID is used to get to the correct data page and slot number on that page.
Once a record is assigned a RID, it does not change until a table reorganization.
Page 473
Page Header Supported page sizes:
3800 -1 3400 4KB, 8KB,
Free space 16KB, 32KB
(usable without page Set on table space creation.
reorganization *) Record 2 Each table space must be
assigned a buffer pool with
Embedded free space Record 1 a matching page size.
(usable after online
page reorganization*)
When a table page is reorganized, embedded free space that is left on the page
after a record is physically deleted is converted to usable free space. RIDs are
redefined based on movement of records on a data page to take advantage of the
usable free space.
Chapter 2. Architecture and processes 19
DB2 supports different page sizes. Use larger page sizes for workloads that tend to
access rows sequentially. For example, sequential access is used for Decision
Support applications or where temporary tables are extensively used. Use smaller
page sizes for workloads that tend to be more random in their access. For example,
random access is used in OLTP environments.
The optimized B-tree implementation has bi-directional pointers on the leaf pages
that allows a single index to support scans in either forward or reverse direction.
Index page are usually split in half except at the high-key page where a 90/10 split
is used. That is, the high ten percent of the index keys are placed on a new page.
This type of index page split is useful for workloads where INSERT requests are
often completed with new high-keys.
Starting in Version 8.1, DB2 uses type-2 indexes. If you migrate from earlier
versions of DB2, both type-1 and type-2 indexes are in use until you reorganize
indexes or perform other actions that convert type-1 indexes to type-2. The index
type determines how deleted keys are physically removed from the index pages.
v For type-1 indexes, keys are removed from the index pages during key deletion
and index pages are freed when the last index key on the page is removed.
v For type-2 indexes, index keys are removed from the page during key deletion
only if there is an X lock on the table. If keys cannot be removed immediately,
they are marked deleted and physically removed later. For more information,
refer to the section that describes type-2 indexes.
Note: Because online defragmentation occurs only when keys are removed from
an index page, in a type-2 index it does not occur if keys are merely marked
deleted, but have not been physically removed from the page.
The INCLUDE clause of the CREATE INDEX statement allows the inclusion of a
specified column or columns on the index leaf pages in addition to the key
columns. This can increase the number of queries that are eligible for index-only
access. However, this can also increase the index space requirements and, possibly,
index maintenance costs if the included columns are updated frequently. The
maintenance cost of updating include columns is less than that of updating key
columns, but more than that of updating columns that do not appear in the index.
Ordering the index B-tree is only done using the key columns and not the included
columns.
K BID
1 4021
C BID
block 0 K BID
2 4022
7 255
8 1488
9 1489
Legend
block 2
reserved for system records
10 1490 FSCR
user records
X reserved
11 1491 U in use
F free
Figure 6. Logical table, record, and index structure for MDC tables
The first block contains special internal records that are used by DB2® to manage
the table, including the free-space control record (FSCR). In subsequent blocks, the
As the name implies, MDC tables cluster data on more than one dimension. Each
dimension is determined by a column or set of columns that you specify in the
ORGANIZE BY DIMENSIONS clause of the CREATE TABLE statement. When you
create an MDC table, the following two kinds of indexes are created automatically:
v A dimension-block index, which contains pointers to each occupied block for a
single dimension.
v A composite block index, which contains all dimension key columns. The
composite block index is used to maintain clustering during insert and update
activity.
The optimizer considers access plans which utilize dimension-block indexes when
it determines the most efficient access plan for a particular query. When queries
have predicates on dimension values, the optimizer can use the dimension block
index to identify, and fetch from, the extents that contain these values. Because
extents are physically contiguous pages on disk, this results in more efficient
performance and minimizes I/O.
In addition, you can create specific RID indexes if analysis of data access plans
indicates that such indexes would improve query performance.
Along with the dimension block indexes and the composite block index, MDC
tables maintain a block map that contains a bitmap that indicates the availability
status of each block. The following attributes are coded in the bitmap list:
v X (reserved): the first block contains only system information for the table.
v U (in use): this block is used and associated with a dimension block index
v L (loaded): this block has been loaded by a current load operation
v C (check constraint): this block is set by the load operation to specify
incremental constraint checking during the load.
v T (refresh table): this block is set by the load operation to specify that AST
maintenance is required.
v F (free): If no other attribute is set, the block is considered free.
Because each block has an entry in the block map file, the file grows as the table
grows. This file is stored as a separate object. In an SMS tablespace it is a new file
type. In a DMS table space, it has a new object descriptor in the object table.
Related concepts:
v “Space requirements for database objects” in the Administration Guide: Planning
v “Designing multidimensional clustering (MDC) tables” in the Administration
Guide: Planning
v “Multidimensional clustering (MDC) table creation, placement, and use” in the
Administration Guide: Planning
v “Insert processing” on page 26
Index structure
The database manager uses a B+ tree structure for index storage. A B+ tree has one
or more levels, as shown in the following diagram, in which RID means row ID:
. .
. .
. .
(‘F’,rid) (‘G’,rid) (‘M’,rid)
LEAF
(‘I’,rid) (‘N’,rid) NODES
(‘K’,rid)
The top level is called the root node. The bottom level consists of leaf nodes in which
the index key values are stored with pointers to the row in the table that contains
the key value. Levels between the root and leaf node levels are called intermediate
nodes.
When it looks for a particular index key value, the index manager searches the
index tree, starting at the root node. The root contains one key for each node at the
next level. The value of each of these keys is the largest existing key value for the
corresponding node at the next level. For example, if an index has three levels as
shown in the figure, then to find an index key value, the index manager searches
the root node for the first key value greater than or equal to the key being looked
for. The root node key points to a specific intermediate node. The index manager
follows this procedure through the intermediate nodes until it finds the leaf node
that contains the index key that it needs.
The figure shows the key being looked for as “I”. The first key in the root node
greater than or equal to “I” is “N”. This points to the middle node at the next
level. The first key in that intermediate node that is greater than or equal to “I” is
“L”. This points to a specific leaf node where the index key for “I” and its
corresponding row ID is found. The row ID identifies the corresponding row in the
base table. The leaf node level can also contain pointers to previous leaf nodes.
These pointers allow the index manager to scan across leaf nodes in either
direction to retrieve a range of values after it finds one value in the range. The
ability to scan in either direction is only possible if the index was created with the
ALLOW REVERSE SCANS clause.
A type-2 index is somewhat larger than a type-1 index and provides features that
minimize next-key locking. The one-byte ridFlag byte stored for each RID on the
leaf page of a type-2 index is used to mark the RID as logically deleted so that it
can be physically removed later. For each variable length column included in the
index, one additional byte stores the actual length of the column value. Type-2
indexes might also be larger than type-1 indexes because some keys might be
marked deleted but not yet physically removed from the index page. After the
DELETE or UPDATE transaction is committed, the keys marked deleted can be
cleaned up.
Related concepts:
v “Advantages and disadvantages of indexes” on page 244
v “Index reorganization” on page 252
v “Online index defragmentation” on page 254
Processes
The following sections provide general descriptions of the DB2 processes.
Log processing
All databases maintain log files that keep records of database changes. There are
two logging strategy choices:
v Circular logging, in which the log records fill the log files and then overwrite
the initial log records in the initial log file. The overwritten log records are not
recoverable.
v Retain log records, in which a log file is archived when it fills with log records.
New log files are made available for log records. Retaining log files enables
roll-forward recovery. Roll-forward recovery reapplies changes to the database
based on completed units of work (transactions) that are recorded in the log.
You can specify that roll-forward recovery is to the end of the logs, or to a
particular point in time before the end of the logs.
Regardless of the logging strategy, all changes to regular data and index pages are
written to the log buffer. The data in the log buffer is written to disk by the logger
process. In the following circumstances, query processing must wait for log data to
be written to disk:
v On COMMIT
v Before the corresponding data pages are written to disk, because DB2® uses
write-ahead logging. The benefit of write-ahead logging is that when a
transaction completes by executing the COMMIT statement, not all of the
changed data and index pages need to be written to disk.
v Before some changes are made to metadata, most of which result from executing
DDL statements
v On writing log records into the log buffer, if the log buffer is full
DB2 manages writing log data to disk in this way in order to minimize processing
delay. In an environment in which many short concurrent transactions occur, most
of the processing delay is caused by COMMIT statements that must wait for log
Changes to large objects (LOBs) and LONG VARCHARs are tracked through
shadow paging. LOB column changes are not logged unless you specify log retain
and the LOB column is defined on the CREATE TABLE statement without the
NOT LOGGED clause. Changes to allocation pages for LONG or LOB data types
are logged like regular data pages.
Related concepts:
v “Update processing” on page 27
v “Client-server processing model” on page 28
Related reference:
v “mincommit - Number of commits to group” on page 403
Insert processing
When SQL statements use INSERT to place new information in a table, an INSERT
search algorithm first searches the Free Space Control Records (FSCRs) to find a
page with enough space. However, even when the FSCR indicates a page has
enough free space, the space may not be usable because it is reserved by an
uncommitted DELETE from another transaction. To ensure that uncommitted free
space is usable, you should COMMIT transactions frequently.
Note: To optimize for INSERT speed at the possible expense of faster table growth,
set the DB2MAXFSCRSEARCH registry variable to a small number. To
optimize for space reuse at the possible expense of INSERT speed, set
DB2MAXFSCRSEARCH to a larger number.
After all FSCRs in the entire table have been searched in this way, the records to be
inserted are appended without additional searching. Searching using the FSCRs is
not done again until space is created somewhere in the table, such as following a
DELETE.
Related concepts:
v “Table and index management for standard tables” on page 18
v “Update processing” on page 27
v “Table and index management for MDC tables” on page 21
Update processing
When an agent updates a page, the database manager uses the following protocol
to minimize the I/O required by the transaction and ensure recoverability.
1. The page to be updated is pinned and latched with an exclusive lock. A log
record is written to the log buffer describing how to redo and undo the change.
As part of this action, a log sequence number (LSN) is obtained and is stored in
the page header of the page being updated.
2. The change is made to the page.
3. The page is unlatched and unfixed.
The page is considered to be “dirty” because changes to the page have not
been written out to disk.
4. The log buffer is updated.
Both the data in the log buffer and the “dirty” data page are forced to disk.
For better performance, these I/Os are delayed until a convenient point, such as
during a lull in the system load, or until necessary to ensure recoverability, or to
limit recovery time. Specifically, a “dirty” page is forced to disk at the following
times:
v When another agent chooses it as a victim.
v When a page cleaner acts on the page as the result of:
– Another agent choosing it as a victim.
– The chngpgs_thresh database configuration parameter percentage value is
exceeded. When this value is exceeded, asynchronous page cleaners wake up
and write changed pages to disk.
If proactive page cleaning is enabled, this value is irrelevant and does not
trigger page cleaning.
– The softmax database configuration parameter percentage value is exceeded.
Once exceeded, asynchronous page cleaners wake up and write changed
pages to disk.
If proactive page cleaning is enabled for the database, and the number of
page cleaners has been properly configured for the database, this value
should never be exceeded.
Related concepts:
v “Log processing” on page 25
v “Client-server processing model” on page 28
Related reference:
v “softmax - Recovery range and soft checkpoint interval” on page 405
v “chngpgs_thresh - Changed pages threshold” on page 370
Note: How DB2® manages client connections depends on whether the connection
concentrator is on or off. The connection concentrator is ON when the
max_connections database manager configuration parameter is set larger than
the max_coordagents configuration parameter.
v If the connection concentrator is OFF, each client application is assigned a
unique EDU called a coordinator agent that coordinates the processing for
that application and communicates with it.
v If the connection concentrator is ON, each coordinator agent can manage
many client connections, one at a time, and might coordinate the other
worker agents to do this work. For Internet applications with many
relatively transient connections, or similar applications with many
relatively small transactions, the connection concentrator improves
performance by allowing many more client applications to be connected.
It also reduces system resource use for each connection.
Each of the circles of the following figure represent engine dispatchable units
(EDUs) which are known as “processes” on UNIX® platforms, and “threads” on
Windows® NT.
Server machine
Coordinator db2agntp
App A
agent
A4
shared memory and semaphores db2agent
A3
A2 db2agntp
logical
agents
A1 db2ipccm
App B
Active
Remote client subagents
db2tcpcm
App B B1 B2 Coordinator db2agntp
agent
B3
B4 db2agent
TCP Idle
B5 subagents
db2wdog db2gds db2agntp
db2sysc db2cart
db2resyn db2dart
db2agent
EDUs per connection EDUs per active database EDUs per request
App A TEST database
db2agntp
db2loggr db2dlock
Fenced processes
This figure shows additional engine dispatchable units (EDUs) that are part of the
server machine environment. Each active database has its own shared pool of
prefetchers (db2pfchr) and page cleaners (db2pclnr), and its own logger (db2loggr)
and deadlock detector (db2dlock).
Fenced user-defined functions (UDFs) and stored procedures, which are not shown
in the figure, are managed to minimize costs associated with their creation and
destruction. The default for the keepfenced database manager configuration
parameter is “YES”, which keeps the stored procedure process available for re-use
at the next stored procedure call.
Note: Unfenced UDFs and stored procedures run directly in an agent’s address
space for better performance. However, because they have unrestricted
access to the agent’s address space, they need to be rigorously tested before
being used.
db2glock db2glock
Catalog node for PROD Catalog node for TEST
Most engine dispatchable units (EDUs) are the same between the single partition
processing model and the multiple partition processing model.
In a multiple partition (or node) environment, one of the partitions is the catalog
node. The catalog keeps all of the information relating to the objects in the
database.
As shown in the figure above, because Application A creates the PROD database
on Node0000, the catalog for the PROD database is created on this node. Similarly,
because Application B creates the TEST database on Node0001, the catalog for the
TEST database is created on this node. You might want to create your databases on
different nodes to balance the extra activity associated with the catalogs for each
database across the nodes in your system environment.
There is also an additional EDU (db2glock) associated with the catalog node for
the database. This EDU controls global deadlocks across the nodes where the
active database is located.
Parts of the database requests from the application are sent by the coordinator
node to subagents at the other partitions; and all results from the other partitions
are consolidated at the coordinator node before being sent back to the application.
The database partition where the CREATE DATABASE command was issued is
called the “catalog node” for the database. It is at this database partition that the
catalog tables are stored. Typically, all user tables are partitioned across a set of
nodes.
Note: Any number of partitions can be configured to run on the same machine.
This is known as a “multiple logical partition”, or “multiple logical node”,
configuration. Such a configuration is very useful on large symmetric
multiprocessor (SMP) machines with very large main memory. In this
environment, communications between partitions can be optimized to use
shared memory and semaphores.
Related concepts:
v “DB2 architecture and process overview” on page 9
v “Log processing” on page 25
v “Update processing” on page 27
v “Memory management” on page 32
v “Connection-concentrator improvements for client connections” on page 259
Memory management
A primary performance tuning task is deciding how to divide the available
memory among the areas within the database. You tune this division of memory
by setting the key configuration parameters described in this section.
All engine dispatchable units (EDUs) in a partition are attached to the Instance
Shared Memory. All EDUs doing work within a database are attached to the
Database Shared Memory of that database. All EDUs working on behalf of a
particular application are attached to an Application Shared Memory region for
that application. This type of shared memory is only allocated if intra- or
inter-partition parallelism is enabled. Finally, each EDU has its own private
memory.
Many different memory areas are contained in database shared memory including:
v Buffer pools
v Lock list
v Database heap – and this includes the log buffer .
v Utility heap
v Package cache
v Catalog cache
Note: Memory can be allocated, freed, and exchanged between different areas
while the database is running. For example, you can decrease the catalog
cache and then increase any given bufferpool by the same amount.
However, before changing the configuration parameter dynamically, you
must be connected to that database. All the memory areas listed above can
be changed dynamically, although the lock list memory area can only be
increased dynamically, and not decreased.
The database manager configuration parameter numdb specifies the number of local
databases that can be concurrently active. The value of the numdb parameter may
impact the total amount of memory allocated.
Agent private memory is allocated for an agent when that agent is created. The
agent private memory contains memory allocations that will be used only by this
specific agent, such as the sort heap and the application heap.
Lock list
Package
cache Buffer pools
Shared
sorts
Database
heap
Extended
buffer pool
(individual segments
attached on demand)
Disks
estore_seg_sz
num_estore_segs (can be > 4Gb)
Extended storage acts as an extended look-aside buffer for the main buffer pools. It
can be much larger than 4 GB. For 32-bit computers with large amounts of main
memory, look-aside buffers can exploit such memory performance improvements.
The extended storage cache is defined in terms of memory segments. For 64-bit
computers, such methods are not needed to access all available memory.
Note, however, that if you use some of the real addressable memory as an
extended storage cache, this memory can no longer be used for other purposes on
The following database configuration parameters influence the amount and size of
the memory available for extended storage:
v num_estore_segs defines the number of extended storage memory segments.
v estore_seg_sz defines the size of each extended memory segment.
Each table space is assigned a buffer pool. An extended storage cache must always
be associated with one or more specific buffer pools. The page size of the extended
storage cache must match the page size of the buffer pool it is associated with.
Related concepts:
v “Organization of memory use” on page 211
v “Database manager shared memory” on page 213
v “Global memory and parameters that control it” on page 216
v “Buffer pool management” on page 220
v “Secondary buffer pools in extended memory on 32-bit platforms” on page 221
v “Guidelines for tuning parameters that affect memory usage” on page 218
v “Connection-concentrator improvements for client connections” on page 259
Related reference:
v “estore_seg_sz - Extended storage memory segment size” on page 373
v “max_coordagents - Maximum number of coordinating agents” on page 379
v “num_estore_segs - Number of extended storage memory segments” on page
373
v “maxagents - Maximum number of agents” on page 380
v “numdb - Maximum number of concurrently active databases including host
and iSeries databases” on page 460
v “max_connections - Maximum number of client connections” on page 379
v “instance_memory - Instance memory” on page 364
v “database_memory - Database shared memory size” on page 338
Concurrency issues
Because many users access and change data in a relational database, the database
manager must be able both to allow users to make these changes and to ensure
that data integrity is preserved. Concurrency refers to the sharing of resources by
multiple interactive users or application programs at the same time. The database
manager controls this access to prevent undesirable effects, such as:
v Lost updates. Two applications, A and B, might both read the same row from
the database and both calculate new values for one of its columns based on the
data these applications read. If A updates the row with its new value and B then
also updates the row, the update performed by A is lost.
v Access to uncommitted data. Application A might update a value in the
database, and application B might read that value before it was committed.
Then, if the value of A is not later committed, but backed out, the calculations
performed by B are based on uncommitted (and presumably invalid) data.
v Nonrepeatable reads. Some applications involve the following sequence of
events: application A reads a row from the database, then goes on to process
other SQL requests. In the meantime, application B either modifies or deletes the
row and commits the change. Later, if application A attempts to read the original
row again, it receives the modified row or discovers that the original row has
been deleted.
v Phantom Read Phenomenon. The phantom read phenomenon occurs when:
1. Your application executes a query that reads a set of rows based on some
search criterion.
2. Another application inserts new data or updates existing data that would
satisfy your application’s query.
3. Your application repeats the query from step 1 (within the same unit of
work).
Some additional (“phantom”) rows are returned as part of the result set that
were not returned when the query was initially executed (step 1).
Note: Declared temporary tables have no concurrency issues because they are
available only to the application that declared them. This type of table only
exists from the time that the application declares it until the application
completes or disconnects.
A DB2 federated system provides location transparency for database objects. For
example, with location transparency if information about tables and views is
moved, references to that information through nicknames can be updated without
changing applications that request the information. When an application accesses
data through nicknames, DB2 relies on the concurrency control protocols of
data-source database managers to ensure isolation levels. Although DB2 tries to
match the requested level of isolation at the data source with a logical equivalent,
results may vary depending on data source capabilities.
Related concepts:
v “Performance impact of isolation levels” on page 40
Related tasks:
v “Specifying the isolation level” on page 43
Related reference:
v “locklist - Maximum storage for lock list” on page 340
v “maxlocks - Maximum percent of lock list before escalation” on page 369
Note: Some host database servers support the no commit isolation level. On other
databases, this isolation level behaves like the uncommitted read isolation
level.
Detailed explanations for each of the isolation levels follows in decreasing order of
performance impact, but in increasing order of care required when accessing and
updating data.
Repeatable Read
Repeatable Read (RR) locks all the rows an application references within a unit of
work. Using Repeatable Read, a SELECT statement issued by an application twice
within the same unit of work in which the cursor was opened, gives the same
result each time. With Repeatable Read, lost updates, access to uncommitted data,
and phantom rows are not possible.
The Repeatable Read application can retrieve and operate on the rows as many
times as needed until the unit of work completes. However, no other applications
With Repeatable Read, every row that is referenced is locked, not just the rows that
are retrieved. Appropriate locking is performed so that another application cannot
insert or update a row that would be added to the list of rows referenced by your
query, if the query was re-executed. This prevents phantom rows from occurring.
For example, if you scan 10 000 rows and apply predicates to them, locks are held
on all 10 000 rows, even though only 10 rows qualify.
Note: The Repeatable Read isolation level ensures that all returned data remains
unchanged until the time the application sees the data, even when temporary
tables or row blocking are used.
Since Repeatable Read may acquire and hold a considerable number of locks, these
locks may exceed the number of locks available as a result of the locklist and
maxlocks configuration parameters. In order to avoid lock escalation, the optimizer
may elect to acquire a single table-level lock immediately for an index scan, if it
believes that lock escalation is very likely to occur. This functions as though the
database manager has issued a LOCK TABLE statement on your behalf. If you do
not want a table-level lock to be obtained ensure that enough locks are available to
the transaction or use the Read Stability isolation level.
Read Stability
Read Stability (RS) locks only those rows that an application retrieves within a unit
of work. It ensures that any qualifying row read during a unit of work is not
changed by other application processes until the unit of work completes, and that
any row changed by another application process is not read until the change is
committed by that process. That is, “nonrepeatable read” behavior is not possible.
Unlike repeatable read, with Read Stability, if your application issues the same
query more than once, you may see additional phantom rows (the phantom read
phenomenon). Recalling the example of scanning 10 000 rows, Read Stability only
locks the rows that qualify. Thus, with Read Stability, only 10 rows are retrieved,
and a lock is held only on those ten rows. Contrast this with Repeatable Read,
where in this example, locks would be held on all 10 000 rows. The locks that are
held can be share, next share, update, or exclusive locks.
Note: The Read Stability isolation level ensures that all returned data remains
unchanged until the time the application sees the data, even when temporary
tables or row blocking are used.
One of the objectives of the Read Stability isolation level is to provide both a high
degree of concurrency as well as a stable view of the data. To assist in achieving
this objective, the optimizer ensures that table level locks are not obtained until
lock escalation occurs.
The Read Stability isolation level is best for applications that include all of the
following:
v Operate in a concurrent environment
v Require qualifying rows to remain stable for the duration of the unit of work
v Do not issue the same query more than once within the unit of work, or do not
require that the query get the same answer when issued more than once in the
same unit of work.
Recalling the example of scanning 10 000 rows, if you use Cursor Stability, you will
only have a lock on the row under your current cursor position. The lock is
removed when you move off that row (unless you update that row).
With Cursor Stability, both nonrepeatable read and the phantom read phenomenon
are possible. Cursor Stability is the default isolation level and should be used when
you want the maximum concurrency while seeing only committed rows from other
applications.
Uncommitted Read
Note: Cursors that are updatable operating under the Uncommitted Read isolation
level will behave as if the isolation level was cursor stability.
When it runs a program using isolation level UR, an application can use isolation
level CS. This happens because the cursors used in the application program are
ambiguous. The ambiguous cursors can be escalated to isolation level CS because
of a BLOCKING option. The default for the BLOCKING option is UNAMBIG. This
means that ambiguous cursors are treated as updatable and the escalation of the
isolation level to CS occurs. To prevent this escalation, you have the following two
choices:
v Modify the cursors in the application program so that they are unambiguous.
Change the SELECT statements to include the FOR READ ONLY clause.
v Leave cursors ambiguous in the application program, but precompile the
program or bind it with the BLOCKING ALL option to allow any ambiguous
cursors to be treated as read-only when the program is run.
As in the example given for Repeatable Read, of scanning 10 000 rows, if you use
Uncommitted Read, you do not acquire any row locks.
With Uncommitted Read, both nonrepeatable read behavior and the phantom read
phenomenon are possible. The Uncommitted Read isolation level is most
The following table summarizes the different isolation levels in terms of their
undesirable effects.
Table 1. Summary of isolation levels
Access to
uncommitted Nonrepeatable Phantom read
Isolation Level data reads phenomenon
Repeatable Read (RR) Not possible Not possible Not possible
Read Stability (RS) Not possible Not possible Possible
Cursor Stability (CS) Not possible Possible Possible
Uncommitted Read (UR) Possible Possible Possible
The table below provides a simple heuristic to help you choose an initial isolation
level for your applications. Consider this table a starting point, and refer to the
previous discussions of the various levels for factors that might make another
isolation level more appropriate.
Table 2. Guidelines for choosing an isolation level
High data stability not
Application Type High data stability required required
Read-write transactions RS CS
Read-only transactions RR or RS UR
Related concepts:
v “Concurrency issues” on page 39
Related tasks:
v “Specifying the isolation level” on page 43
The isolation level can be specified in several different ways. The following
heuristics are used in determining which isolation level will be used in compiling
an SQL statement:
Dynamic SQL:
v If an isolation clause is specified in the statement, then the value of that clause is
used.
v If no isolation clause is specifed in the statement, and a SET CURRENT
ISOLATION statement has been issued within the current session, then the value
of the CURRENT ISOLATION special register is used.
v If no isolation clause is specifed in the statement, and no SET CURRENT
ISOLATION statement has been issued within the current session, then the
isolation level used is the one specified for the package at the time when the
package was bound to the database.
Note: Many commercially written applications provide a method for choosing the
isolation level. Refer to the application documentation for information.
Procedure:
where XXXXXXXX is the name of the package and YYYYYYYY is the schema
name of the package. Both of these names must be in all capital letters.
2. On database servers that support REXX:
When a database is created, multiple bind files that support the different
isolation levels for SQL in REXX are bound to the database. Other
command-line processor packages are also bound to the database when a
database is created.
REXX and the command line processor connect to a database using a default
isolation level of cursor stability. Changing to a different isolation level does not
change the connection state. It must be executed in the CONNECTABLE AND
UNCONNECTED state or in the IMPLICITLY CONNECTABLE state.
To verify the isolation level in use by a REXX application, check the value of
the SQLISL REXX variable. The value is updated every time the CHANGE
SQLISL command is executed.
Note: JDBC and SQLJ are implemented with CLI on DB2, which means the
[Link] settings might affect what is written and run using JDBC and
SQLJ.
Use the setTransactionIsolation method in the [Link] interface connection.
In SQLJ, you run the db2profc SQLJ optimizer to create a package. The options
that you can specify for this package include its isolation level.
6. For dynamic SQL within the current session:
Use the SET CURRENT ISOLATION statement to set the isolation level for
dynamic SQL issued within a session. Issuing this statement sets the CURRENT
ISOLATION special register to a value that specifies the level of isolation for
any dynamic SQL statements issued within the current session. Once set, the
CURRENT ISOLATION special register provides the isolation level for any
subsequent dynamic SQL statement compiled within the session, regardless of
the package issuing the statement. This isolation level will apply until the
session is ended or until a SET CURRENT ISOLATION statement is issued
with the RESET option.
Related concepts:
v “Concurrency issues” on page 39
Related reference:
v “SQLSetConnectAttr function (CLI) - Set connection attributes” in the CLI Guide
and Reference, Volume 2
v “CONNECT (Type 1) statement” in the SQL Reference, Volume 2
v “Statement attributes (CLI) list” in the CLI Guide and Reference, Volume 2
Although most locking occurs on tables, when a buffer pool is created, altered, or
dropped, a buffer pool lock is set. The mode used with this lock is EXCLUSIVE
(X). You may encounter this lock when a snapshot is taken using the Command
Line Processor (CLP). When viewing the snapshot, you will see that the lock name
used is the identifier (ID) of the buffer pool itself.
In general, record-level locking is used unless one of the following is the case:
v The isolation level chosen is uncommitted read (UR).
v The isolation level chosen is repeatable read (RR) and the access plan requires a
scan with no predicates.
v The table LOCKSIZE attribute is “TABLE”.
v The lock list fills, causing escalation.
A lock escalation occurs when the number of locks held on rows and tables in the
database equals the percentage of the locklist specified by the maxlocks database
configuration parameter. Lock escalation might not affect the table that acquires the
lock that triggers escalation. To reduce the number of locks to about half the
number held when lock escalation begins, the database manager begins converting
many small row locks to table locks for all active tables, beginning with any locks
on large object (LOB) or long VARCHAR elements. An exclusive lock escalation is a
lock escalation in which the table lock acquired is an exclusive lock. Lock
escalations reduce concurrency. Conditions that might cause lock escalations
should be avoided.
The duration of row locking varies with the isolation level being used:
v UR scans: No row locks are held unless row data is changing.
v CS scans: Row locks are only held while the cursor is positioned on the row.
v RS scans: Only qualifying row locks are held for the duration of the transaction.
v RR scans: All row locks are held for the duration of the transaction.
Related concepts:
v “Lock attributes” on page 47
v “Locks and performance” on page 49
v “Guidelines for locking” on page 53
Related reference:
v “locklist - Maximum storage for lock list” on page 340
v “dlchktime - Time interval for checking deadlock” on page 367
v “diaglevel - Diagnostic error capture level” on page 451
v “locktimeout - Lock timeout” on page 368
v “Lock type compatibility” on page 59
v “Lock modes and access paths for standard tables” on page 60
v “Lock modes for table and RID index scans of MDC tables” on page 62
v “Locking for block index scans for MDC tables” on page 65
Lock attributes
Database manager locks have the following basic attributes:
Mode The type of access allowed for the lock owner as well as the type of access
permitted for concurrent users of the locked object. It is sometimes referred
to as the state of the lock.
Object
The resource being locked. The only type of object that you can lock
explicitly is a table. The database manager also imposes locks on other
types of resources, such as rows, tables, and table spaces. For
multidimensional clustering (MDC) tables, block locks can also be
imposed. The object being locked determines the granularity of the lock.
The following table shows the modes and their effects in order of increasing
control over resources. For detailed information about locks at various levels, refer
to the lock-mode reference tables.
Table 3. Lock Mode Summary
Applicable Object
Lock Mode Type Description
IN (Intent None) Table spaces, blocks, The lock owner can read any data in the object, including
tables uncommitted data, but cannot update any of it. Other concurrent
applications can read or update the table.
IS (Intent Share) Table spaces, blocks, The lock owner can read data in the locked table, but cannot update
tables this data. Other applications can read or update the table.
NS (Next Key Share) Rows The lock owner and all concurrent applications can read, but not
update, the locked row. This lock is acquired on rows of a table,
instead of an S lock, where the isolation level of the application is
either RS or CS. NS lock mode is not used for next-key locking. It is
used instead of S mode during CS and RS scans to minimize the
impact of next-key locking on these scans.
S (Share) Rows, blocks, tables The lock owner and all concurrent applications can read, but not
update, the locked data.
IX (Intent Exclusive) Table spaces, blocks, The lock owner and concurrent applications can read and update
tables data. Other concurrent applications can both read and update the
table.
SIX (Share with Tables, blocks The lock owner can read and update data. Other concurrent
Intent Exclusive) applications can read the table.
U (Update) Rows, blocks, tables The lock owner can update data. Other units of work can read the
data in the locked object, but cannot attempt to update it.
NW (Next Key Weak Rows When a row is inserted into an index, an NW lock is acquired on
Exclusive) the next row. For type 2 indexes, this occurs only if the next row is
currently locked by an RR scan. The lock owner can read but not
update the locked row. This lock mode is similar to an X lock,
except that it is also compatible with W and NS locks.
X (Exclusive) Rows, blocks, tables, The lock owner can both read and update data in the locked object.
buffer pools Only uncommitted read applications can access the locked object.
W (Weak Exclusive) Rows This lock is acquired on the row when a row is inserted into a table
that does not have type-2 indexes defined. The lock owner can
change the locked row. To determine if a duplicate value has been
committed when a duplicate value is found, this lock is also used
during insertion into a unique index. This lock is similar to an X
lock except that it is compatible with the NW lock. Only
uncommitted read applications can access the locked row.
Z (Super Exclusive) Table spaces, tables This lock is acquired on a table in certain conditions, such as when
the table is altered or dropped, an index on the table is created or
dropped, or for some types of table reorganization. No other
concurrent application can read or update the table.
Related concepts:
v “Locks and concurrency control” on page 46
v “Locks and performance” on page 49
If one application holds a lock on a database object, another application might not
be able to access that object. For this reason, row-level locks are better for
maximum concurrency than table-level locks. However, locks require storage and
processing time, so a single table lock minimizes lock overhead.
The LOCKSIZE clause of the ALTER TABLE statement specifies the scope
(granularity) of locks at either row or table level. By default, row locks are used.
Only S (Shared) and X (Exclusive) locks are requested by these defined table locks.
The ALTER TABLE statement LOCKSIZE ROW clause does not prevent normal
lock escalation from occurring.
The ALTER TABLE statement specifies locks globally, affecting all applications and
users that access that table. Individual applications might use the LOCK TABLE
statement to specify table locks at an application level instead.
Lock compatibility
Assume that application A holds a lock on a table that application B also wants to
access. The database manager requests, on behalf of application B, a lock of some
particular mode. If the mode of the lock held by A permits the lock requested by
B, the two locks (or modes) are said to be compatible.
Lock conversion
Changing the mode of a lock already held is called a conversion. Lock conversion
occurs when a process accesses a data object on which it already holds a lock, and
the access mode requires a more restrictive lock than the one already held. A
process can hold only one lock on a data object at any time, although it can
request a lock many times on the same data object indirectly through a query.
Some lock modes apply only to tables, others only to rows or blocks. For rows or
blocks, conversion usually occurs if an X is needed and an S or U (Update) lock is
held.
IX (Intent Exclusive) and S (Shared) locks are special cases with regard to lock
conversion, however. Neither S nor IX is considered to be more restrictive than the
other, so if one of these is held and the other is required, the resulting conversion
is to a SIX (Share with Intent Exclusive) lock. All other conversions result in the
requested lock mode becoming the mode of the lock held if the requested mode is
more restrictive.
A dual conversion might also occur when a query updates a row. If the row is read
through an index access and locked as S, the table that contains the row has a
covering intention lock. But if the lock type is IS instead of IX, if the row is
subsequently changed the table lock is converted to an IX and the row to an X.
Lock Escalation
Lock escalation is an internal mechanism that reduces the number of locks held. In
a single table, locks are escalated to a table lock from many row locks, or for
multi-dimensional clustering (MDC) tables, from many row or block locks. Lock
escalation occurs when applications hold too many locks of any type. Lock
escalation can occur for a specific database agent if the agent exceeds its allocation
of the lock list. Such escalation is handled internally; the only externally detectable
result might be a reduction in concurrent access on one or more tables.
Occasionally, the process receiving the internal escalation request holds few or no
row locks on any table, but locks are escalated because one or more processes hold
Note: Lock escalation might also cause deadlocks. For example, suppose a
read-only application and an update application are both accessing the same
table. If the update application has exclusive locks on many rows on the
table, the database manager might try to escalate the locks on this table to
an exclusive table lock. However, the table lock held by the read-only
application will cause the exclusive lock escalation request to wait. If the
read-only application requires a row lock on a row already locked by the
update application, this creates a deadlock. To avoid this kind of problem,
either code the update application to lock the table exclusively when it starts
or increase the size of the lock list.
Setting this parameter helps avoid global deadlocks, especially in distributed unit
of work (DUOW) applications. If the time that the lock request is pending is
greater than the locktimeout value, the requesting application receives an error and
its transaction is rolled back. For example, if program1 tries to acquire a lock which
is already held by program2, program1 returns SQLCODE -911 (SQLSTATE 40001)
with reason code 68 if the timeout period expires. The default value for locktimeout
is -1, which turns off lock timeout detection.
| Note: For table, row, and MDC block locks, an application can override the
| database level locktimeout setting by using SET CURRENT LOCK TIMEOUT.
Deadlocks
Contention for locks can result in deadlocks. For example, suppose that Process 1
locks table A in X (exclusive) mode and Process 2 locks table B in X mode. If
Process 1 then tries to lock table B in X mode and Process 2 tries to lock table A in
X mode, the processes are in a deadlock. In a deadlock, both processes are
suspended until their second lock request is granted, but neither request is granted
until one of the processes performs a commit or rollback. This state continues
indefinitely until an external agent activates one of the processes and forces it to
perform a rollback.
If it finds a deadlock, the deadlock detector selects one deadlocked process as the
victim process to roll back. The victim process is awakened, and returns SQLCODE
-911 (SQLSTATE 40001), with reason code 2, to the calling application. The
database manager rolls back the selected process automatically. When the rollback
is complete, the locks that belonged to the victim process are released, and the
other processes involved in the deadlock can continue.
To ensure good performance, select the proper interval for the deadlock detector.
An interval that is too short causes unnecessary overhead, and an interval that is
too long allows a deadlock to delay a process for an unacceptable amount of time.
For example, a wake-up interval of 5 minutes allows a deadlock to exist for almost
5 minutes, which can seem like a long time for short transaction processing.
Balance® the possible delays in resolving deadlocks with the overhead of detecting
them.
| A different problem occurs when an application with more than one independent
| process that accesses the database is structured to make deadlocks likely. An
| example is an application in which several processes access the same table for
| reads and then writes. If the processes do read-only SQL queries at first and then
| do SQL updates on the same table, the chance of deadlocks increases because of
| potential contention between the processes for the same data. For instance, if two
| processes read the table, and then update the table, process A might try to get an X
| lock on a row on which process B has an S lock, and vice versa. To avoid such
| deadlocks, applications that access data with the intention of modifying it should
| do one of the following:
| v Use the FOR UPDATE OF clause when performing a select. This clause ensures
| that a U lock is imposed when process A attempts to read the data. Row
| blocking, however, is disabled.
| v Use the WITH RR USE AND KEEP UPDATE LOCKS or the WITH RS USE AND
| KEEP UPDATE LOCKS clause when performing the query. Either clause ensures
| that a U lock is imposed when process A attempts to read the data, and allows
| row blocking.
Note: You might consider defining a monitor that records when deadlocks occur.
Use the SQL statement CREATE EVENT to create a monitor.
To limit the amount of disk space that this event monitor consumes, the
event monitor deactivates, and a message is written to the administration
To log more information about deadlocks, set the database manager configuration
parameter notifylevel to four. The administration notification log stores information
that includes the object, the lock mode, and the application holding the lock on the
object. The current dynamic SQL statement or static package name might also be
logged. The dynamic SQL statement is logged only if notifylevel is four.
Related concepts:
v “Locks and concurrency control” on page 46
v “Deadlocks between applications” on page 11
Related tasks:
v “Correcting lock escalation problems” on page 55
Related reference:
v “diaglevel - Diagnostic error capture level” on page 451
v “locktimeout - Lock timeout” on page 368
Related concepts:
v “Locks and concurrency control” on page 46
v “Lock attributes” on page 47
v “Locks and performance” on page 49
v “Factors that affect locking” on page 68
In a well designed database, lock escalation rarely occurs. If lock escalation reduces
concurrency to an unacceptable level, however, you need to analyze the problem
and decide how to solve it.
Prerequisites:
Ensure that lock escalation information is recorded. Set the database manager
configuration parameter notifylevel to 3, which is the default, or to 4. At notifylevel
of 2, only the error SQLCODE is reported. At notifylevel of 3 or 4, when lock
escalation fails, information is recorded for the error SQLCODE and the table for
which the escalation failed. The current SQL statement is logged only if it is a
currently executing, dynamic SQL statement and notifylevelis set to 4.
Procedure:
Follow these general steps to diagnose the cause of unacceptable lock escalations
and apply a remedy:
1. Analyze in the administration notification log on all tables for which locks are
escalated. This log file includes the following information:
v The number of locks currently held.
v The number of locks needed before lock escalation is completed.
v The table identifier information and table name of each table being escalated.
v The number of non-table locks currently held.
v The new table level lock to be acquired as part of the escalation. Usually, an
“S,” or Share lock, or an “X,” or eXclusive lock is acquired.
v The internal return code of the result of the acquisition of the new table lock
level.
2. Use the information in administration notification log to decide how to resolve
the escalation problem. Consider the following possibilities:
Related reference:
v “maxlocks - Maximum percent of lock list before escalation” on page 369
v “diaglevel - Diagnostic error capture level” on page 451
With this variable enabled, predicate evaluation can occur on uncommitted data.
This means that a row that contains an uncommitted update may not satisfy the
query, whereas if the predicate evaluation waited until the updated transaction
completed, the row may satisfy the query. Additionally, uncommitted deleted rows
are skipped during table scans. DB2 will skip deleted keys in type-2 index scans if
the DB2_SKIPDELETED registry variable is enabled.
These registry variable settings apply at compile time for dynamic SQL and at bind
time for static SQL. This means that even if the registry variable is enabled at
runtime, the lock avoidance strategy is not employed unless
DB2_EVALUNCOMMITTED was enabled at bind time. If the registry variable is
enabled at bind time but not enabled at runtime, the lock avoidance strategy is still
in effect. For static SQL, if a package is rebound, the registry variable setting at
bind time is the setting that applies. An implicit rebind of static SQL will use the
current setting of the DB2_EVALUNCOMMITTED.
Example
The following example provides a comparison of the default locking behavior and
the new evaluate uncommitted behavior.
The table below is the ORG table from the SAMPLE database.
DEPTNUMB DEPTNAME MANAGER DIVISION LOCATION
-------- -------------- ------- ---------- -------------
10 Head Office 160 Corporate New York
15 New England 50 Eastern Boston
20 Mid Atlantic 10 Eastern Washington
38 South Atlantic 30 Eastern Atlanta
42 Great Lakes 100 Midwest Chicago
51 Plains 140 Midwest Dallas
66 Pacific 270 Western San Francisco
84 Mountain 290 Western Denver
The following transactions are acting on this table, with the default Cursor Stability
(CS) isolation level.
Table 8. Transactions on the ORG table with the CS isolation level
SESSION 1 SESSION 2
The uncommitted UPDATE in Session 1 holds an exclusive record lock on the first
row in the table, prohibiting the SELECT query in Session 2 from returning even
though the row being updated in Session 1 does not currently satisfy the query in
Session 2. This is because the CS isolation level dictates that any row accessed by a
query must be locked while the cursor is positioned on that row. Session 2 cannot
obtain a lock on the first row until Session 1 releases its lock.
When scanning the table, the lock-wait in Session 2 can be avoided using the
evaluate uncommitted feature which first evaluates the predicate and then locks
the row for a true predicate evaluation. As such, the query in Session 2 would not
attempt to lock the first row in the table thereby increasing application
concurrency. Note that this would also mean that predicate evaluation in Session 2
would occur with respect to the uncommitted value of deptnumb=5 in Session 1.
The query in Session 2 would omit the first row in its result set despite the fact
that a rollback of the update in Session 1 would satisfy the query in Session 2.
If the order of operations were reversed, concurrency could still be improved with
evaluate uncommitted. Under default locking behavior, Session 2 would first
acquire a row lock prohibiting the searched UPDATE in Session 1 from executing
even though the UPDATE in Session 1 would not change the row locked by the
query of Session 2. If the searched UPDATE in Session 1 first attempted to examine
rows and then only lock them if they qualified, the query in Session 1 would be
non-blocking.
Restrictions
The following external restrictions apply to this new functionality:
v The registry variable DB2_EVALUNCOMMITTED must be enabled.
v The isolation level must be CS or RS.
v Row locking is to occur.
v SARGable evaluation predicates exist.
v Evaluation uncommitted is not applicable to scans on the catalog tables.
v For MDC tables, block locking can be deferred for an index scan; however, block
locking will not be deferred for table scans.
v Deferred locking will not occur on a table which is executing an inplace table
reorg.
v Deferred locking will not occur for an index scan where the index is type-1.
v For Iscan-Fetch plans, row locking is not deferred to the data access but rather
the row is locked during index access before moving to the row in the table.
v Deleted rows are unconditionally skipped for table scans while deleted type-2
index keys are only skipped if the registry variable DB2_SKIPDELETED is
enabled.
Related concepts:
v “Locks and concurrency control” on page 46
v “Lock attributes” on page 47
v “Locks and performance” on page 49
Related reference:
v “Lock modes and access paths for standard tables” on page 60
v “Locking for block index scans for MDC tables” on page 65
The following tables list the types of locks obtained for standard tables at each
level for different access plans. Each entry is made up of two parts: table lock and
row lock. A dash indicates that a particular level of locking is not done.
Notes:
1. In a multi-dimensional clustering (MDC) environment, an additional lock level,
BLOCK, is used.
| 2. Lock modes can be changed explicitly with the lock-request-clause of a select
| statement.
Table 10. Lock Modes for Table Scans
Isolation Read-only and Cursored operation Searched update or
Level ambiguous scans delete
Scan Where Scan Update or
current of delete
Access Method: Table scan with no predicates
RR S/- U/- SIX/X X/- X/-
RS IS/NS IX/U IX/X IX/X IX/X
CS IS/NS IX/U IX/X IX/X IX/X
UR IN/- IX/U IX/X IX/X IX/X
Access Method: Table Scan with predicates
RR S/- U/- SIX/X U/- SIX/X
RS IS/NS IX/U IX/X IX/U IX/X
CS IS/NS IX/U IX/X IX/U IX/X
UR IN/- IX/U IX/X IX/U IX/X
Note: At UR isolation level with IN lock for type-1 indexes or if there are
predicates on include columns in the index, the isolation level is upgraded
to CS and the locks to an IS table lock and NS row locks.
Table 11. Lock Modes for RID Index Scans
Isolation Read-only and Cursored operations Searched update or
Level ambiguous scans delete
Scan Where Scan Update or
current of Delete
Access Method: RID index scan with no predicates
RR S/- IX/S IX/X X/- X/-
RS IS/NS IX/U IX/X IX/X IX/X
CS IS/NS IX/U IX/X IX/X IX/X
UR IN/- IX/U IX/X IX/X IX/X
Access Method: RID index scan with a single qualifying row
RR IS/S IX/U IX/X IX/X IX/X
RS IS/NS IX/U IX/X IX/X IX/X
The following table shows the lock modes for cases in which reading of the data
pages is deferred to allow the list of rows to be:
v Further qualified using multiple indexes
v Sorted for efficient prefetching
Table 12. Lock modes for index scans used for deferred data page access
Isolation Read-only and Cursored operations Searched update or
Level ambiguous scans delete
Scan Where Scan Update or
current of delete
Access Method: RID index scan with no predicates
RR IS/S IX/S X/-
RS IN/- IN/- IN/-
CS IN/- IN/- IN/-
UR IN/- IN/- IN/-
Access Method: Deferred Data Page Access, after a RID index scan with no
predicates
RR IN/- IX/S IX/X X/- X/-
RS IS/NS IX/U IX/X IX/X IX/X
CS IS/NS IX/U IX/X IX/X IX/X
UR IN/- IX/U IX/X IX/X IX/X
Access Method: RID index scan with predicates (sargs, resids)
RR IS/S IX/S IX/S
RS IN/- IN/- IN/-
CS IN/- IN/- IN/-
UR IN/- IN/- IN/-
Related concepts:
v “Lock attributes” on page 47
v “Locks and performance” on page 49
Related reference:
v “Lock type compatibility” on page 59
v “Lock modes for table and RID index scans of MDC tables” on page 62
v “Locking for block index scans for MDC tables” on page 65
Lock modes for table and RID index scans of MDC tables
In a multi-dimensional clustering (MDC) environment, an additional lock level,
BLOCK, is used. The following tables list the types of locks obtained at each level
for different access plans. Each entry is made up of three parts: table lock, block
lock, and row lock. A dash indicates that a particular level of locking is not used.
The following two tables show lock modes for RID indexes on MDC tables.
Table 14. Lock Modes for RID Index Scans
Isolation Read-only and Cursored operations Searched update or
Level ambiguous scans delete
Scan Where Scan Update
current of Delete
Access Method: RID index scan with no predicates
RR S/-/- IX/IX/S IX/IX/X X/-/- X/-/-
RS IS/IS/NS IX/IX/U IX/IX/X X/X/X X/X/X
CS IS/IS/NS IX/IX/U IX/IX/X X/X/X X/X/X
UR IN/IN/- IX/IX/U IX/IX/X X/X/X X/X/X
Access Method: RID index scan with single qualifying row
RR IS/IS/S IX/IX/U IX/IX/X X/X/X X/X/X
RS IS/IS/NS IX/IX/U IX/IX/X X/X/X X/X/X
CS IS/IS/NS IX/IX/U IX/IX/X X/X/X X/X/X
UR IN/IN/- IX/IX/U IX/IX/X X/X/X X/X/X
Access Method: RID index scan with start and stop predicates only
RR IS/IS/S IX/IX/S IX/IX/X IX/IX/X IX/IX/X
RS IS/IS/NS IX/IX/U IX/IX/X IX/IX/X IX/IX/X
CS IS/IS/NS IX/IX/U IX/IX/X IX/IX/X IX/IX/X
UR IN/IN/- IX/IX/U IX/IX/X IX/IX/X IX/IX/X
Access Method: Index scan with index predicates only
Note: In the following table, which shows lock modes for RID index scans used
for deferred data-page access, at UR isolation level with IN lock for type-1
indexes or if there are predicates on include columns in the index, the
isolation level is upgraded to CS and the locks are upgraded to an IS table
lock, an IS block lock, and NS row locks.
Table 15. Lock modes for RID index scans used for deferred data-page access
Isolation Read-only and Cursored operations Searched update or
Level ambiguous scans delete
Scan Where Scan Update
current of Delete
Access Method: RID index scan with no predicates
RR IS/S/S IX/IX/S X/-/-
RS IN/IN/- IN/IN/- IN/IN/-
CS IN/IN/- IN/IN/- IN/IN/-
UR IN/IN/- IN/IN/- IN/IN/-
Access Method: Deferred data-page access after a RID index scan with no
predicates
RR IN/IN/- IX/IX/S IX/IX/X X/-/- X/-/-
RS IS/IS/NS IX/IX/U IX/IX/X IX/IX/X IX/IX/X
CS IS/IS/NS IX/IX/U IX/IX/X IX/IX/X IX/IX/X
UR IN/IN/- IX/IX/U IX/IX/X IX/IX/X IX/IX/X
Access Method: RID index scan with predicates (sargs, resids)
RR IS/S/- IX/IX/S IX/IX/S
RS IN/IN/- IN/IN/- IN/IN/-
CS IN/IN/- IN/IN/- IN/IN/-
UR IN/IN/- IN/IN/- IN/IN/-
Access Method: Deferred data-page access after a RID index scan with
predicates (sargs, resids)
RR IN/IN/- IX/IX/S IX/IX/X IX/IX/S IX/IX/X
Related concepts:
v “Locks and concurrency control” on page 46
v “Lock attributes” on page 47
v “Locks and performance” on page 49
Related reference:
v “Lock type compatibility” on page 59
v “Lock modes and access paths for standard tables” on page 60
v “Locking for block index scans for MDC tables” on page 65
The following table lists lock modes for block index scans used for deferred
data-page access:
Table 17. Lock modes for block index scans used for deferred data-page access
Isolation Read-only and Cursored operations Searched update or
Level ambiguous scans delete
Scan Where Scan Update
current of Delete
Access Method: Block index scan with no predicates
RR IS/S/-- IX/IX/S X/--/--
RS IN/IN/-- IN/IN/-- IN/IN/--
CS IN/IN/-- IN/IN/-- IN/IN/--
UR IN/IN/-- IN/IN/-- IN/IN/--
Access Method: Deferred data-page access after a block index scan with no
predicates
RR IN/IN/-- IX/IX/S IX/IX/X X/--/-- X/--/--
RS IS/IS/NS IX/IX/U IX/IX/X X/X/-- X/X/--
CS IS/IS/NS IX/IX/U IX/IX/X X/X/-- X/X/--
UR IN/IN/-- IX/IX/U IX/IX/X X/X/-- X/X/--
Access Method: Block index scan with dimension predicates only
RR IS/S/-- IX/IX/-- IX/S/--
Related concepts:
v “Locks and performance” on page 49
Related reference:
v “Lock type compatibility” on page 59
v “Lock modes and access paths for standard tables” on page 60
v “Lock modes for table and RID index scans of MDC tables” on page 62
Related concepts:
v “Locks and concurrency control” on page 46
v “Lock attributes” on page 47
v “Locks and performance” on page 49
v “Guidelines for locking” on page 53
v “Index cleanup and maintenance” on page 251
v “Locks and types of application processing” on page 68
v “Locks and data-access methods” on page 69
v “Index types and next-key locking” on page 70
A statement that inserts, updates or deletes data in a target table, based on the
result from a sub-select statement, does two types of processing. The rules for
read-only processing determine the locks for the tables returned in the sub-select
statement. The rules for change processing determine the locks for the target table.
Related tasks:
v “Correcting lock escalation problems” on page 55
Related reference:
v “Lock type compatibility” on page 59
If an index is not used, the entire table must be scanned in sequence to find the
selected rows, and may thus acquire a single table level lock (S). For example, if
there is no index on the column SEX, a table scan might be used to select all male
employees with a a statement that contains the following SELECT clause:
SELECT *
FROM EMPLOYEE
WHERE SEX = 'M';
Note: Cursor controlled processing uses the lock mode of the underlying cursor
until the application finds a row to update or delete. For this type of
processing, no matter what the lock mode of a cursor, an exclusive lock is
always obtained to perform the update or delete.
Reference tables provide detailed information about which locks are obtained for
what kind of access plan.
Deferred access of the data pages implies that access to the row occurs in two
steps, which results in more complex locking scenarios. The timing of lock
aquisition and the persistence of the locks depend on the isolation level. Because
the Repeatable Read isolation level retains all locks until the end of the transaction,
the locks acquired in the first step are held and there is no need to acquire further
Related concepts:
v “Locks and concurrency control” on page 46
v “Lock attributes” on page 47
v “Locks and performance” on page 49
v “Guidelines for locking” on page 53
v “Locks and types of application processing” on page 68
v “Index types and next-key locking” on page 70
Related tasks:
v “Correcting lock escalation problems” on page 55
Related reference:
v “Lock type compatibility” on page 59
v “Lock modes and access paths for standard tables” on page 60
v “Lock modes for table and RID index scans of MDC tables” on page 62
v “Locking for block index scans for MDC tables” on page 65
Related concepts:
v “Advantages and disadvantages of indexes” on page 244
v “Index performance tips” on page 248
v “Index structure” on page 23
v “Index reorganization” on page 252
v “Online index defragmentation” on page 254
v “Index cleanup and maintenance” on page 251
Optimization factors
This section describes the factors to consider when you specify the optimization
class for queries.
Note: In a federated database query, the optimization class does not apply to the
remote optimizer.
Setting the optimization class can provide some of the advantages of explicitly
specifying optimization techniques, particularly for the following reasons:
v To manage very small databases or very simple dynamic queries
v To accommodate memory limitations at compile time on your database server
v To reduce the query compilation time, such as PREPARE.
Query optimization classes 1, 2, 3, 5, and 7 are all suitable for general-purpose use.
Consider class 0 only if you require further reductions in query compilation time
and you know that the SQL statements are extremely simple.
Tip: To analyze queries that run a long time, run the query with db2batch to find
out how much time is spent in compilation and how much is spent in
[Link] compilation requires more time, reduce the optimization class. If
execution requires more time, consider a higher optimization class.
When you select an optimization class, consider the following general guidelines:
v Start by using the default query optimization class, class 5.
v To use a class other than the default, try class 1, 2 or 3 first. Classes 0, 1, and 2
use the Greedy join enumeration algorithm.
v Use optimization class 1 or 2 if you have many tables with many of the join
predicates that are on the same column, and if compilation time is a concern.
v Use a low optimization class (0 or 1) for queries having very short run-times of
less than one second. Such queries tend to have the following characteristics:
– Access to a single or only a few tables
– Fetch a single or only a few rows
– Use fully qualified, unique indexes.
Online transaction processing (OLTP) transactions are good examples of this
kind of SQL.
v Use a higher optimization class (3, 5, or 7) for longer running queries that take
more than 30 seconds.
v Classes 3 and above use the Dynamic Programming join enumeration
algorithm. This algorithm considers many more alternative plans, and might
incur significantly more compilation time than classes 0, 1, and 2, especially as
the number of tables increases.
Complex queries might require different amounts of optimization to select the best
access plan. Consider using higher optimization classes for queries that have the
following characteristics:
v Access to large tables
v A large number of predicates
v Many subqueries
v Many joins
v Many set operators, such as UNION and INTERSECT
v Many qualifying rows
v GROUP BY and HAVING operations
v Nested table expressions
v A large number of views.
Use higher query optimization classes for SQL that was produced by a query
generator. Many query generators create inefficient SQL. Poorly written queries,
including those produced by a query generator, require additional optimization to
select a good access plan. Using query optimization class 2 and higher can improve
such SQL queries.
Related concepts:
v “Configuration parameters that affect query optimization” on page 136
v “Benchmark testing” on page 303
v “Optimization strategies for intra-partition parallelism” on page 173
v “Optimization strategies for MDC tables” on page 175
Related tasks:
v “Setting the optimization class” on page 76
Related reference:
v “Optimization classes” on page 73
Optimization classes
You can specify one of the following optimizer classes when you compile an SQL
query:
0- This class directs the optimizer to use minimal optimization to generate an
access plan. This optimization class has the following characteristics:
v Non-uniform distribution statistics are not considered by the optimizer.
v Only basic query rewrite rules are applied.
v Greedy join enumeration occurs.
v Only nested loop join and index scan access methods are enabled.
v List prefetch and index ANDing are not used in generated access
methods.
v The star-join strategy is not considered.
This class should only be used in circumstances that require the the lowest
possible query compilation overhead. Query optimization class 0 is
Related concepts:
v “Optimization class guidelines” on page 72
v “Optimization strategies for intra-partition parallelism” on page 173
v “Remote SQL generation and global optimization in federated databases” on
page 184
v “Optimization strategies for MDC tables” on page 175
If you think that a query that might benefit from additional optimization, but you
are not sure, or you are concerned about compilation time and resource usage, you
might perform some benchmark testing.
Procedure:
Note: After the initial compilation, dynamic SQL statements are recompiled
when a change to the environment requires it. If the environment does
not change after an SQL statement is cached, it does not need to be
compiled again because subsequent PREPARE statements re-use the
cached statement.
v For static SQL statements, compare the statement run times.
Although you might also be interested in the compile time of static SQL, the
total compile and run time for the statement is difficult to assess in any
meaningful context. Comparing the total time does not recognize the fact that
a static SQL statement can be run many times for each time it is bound and
that it is generally not bound during run time.
2. Specify the optimization class as follows:
v Dynamic SQL statements use the optimization class specified by the
CURRENT QUERY OPTIMIZATION special register that you set with the
SQL statement SET. For example, the following statement sets the
optimization class to 1:
SET CURRENT QUERY OPTIMIZATION = 1
To ensure that a dynamic SQL statement always uses the same optimization
class, you might include a SET statement in the application program.
Related concepts:
v “Optimization class guidelines” on page 72
Related reference:
v “Optimization classes” on page 73
Tuning applications
This section provides guidelines for tuning the queries that applications execute.
To improve performance for such applications, you can modify the SELECT
statement in the following ways:
v Use the FOR UPDATE clause to specify the columns that could be updated by a
subsequent positioned UPDATE statement.
v Use the FOR READ/FETCH ONLY clause to make the returned columns read
only.
v Use the OPTIMIZE FOR n ROWS clause to give priority to retrieving the first n
rows in the full result set.
v Use the FETCH FIRST n ROWS ONLY clause to retrieve only a specified number
of rows.
v Use the DECLARE CURSOR WITH HOLD statement to retrieve rows one at a
time.
| Note: Row blocking is affected if you use the FOR UPDATE, FETCH FIRST n
| ROWS ONLY, or the OPTIMIZE FOR n ROWS clause or if you declare your
| cursor as SCROLLing.
The FOR READ ONLY clause or FOR FETCH ONLY clause ensures that read-only
results are returned. Because the result table from a SELECT on a view defined as
read-only is also read only, this clause is permitted but has no effect.
For result tables where updates and deletes are allowed, specifying FOR READ
ONLY may improve the performance of FETCH operations if the database
manager can retrieve blocks of data instead of exclusive locks. Do not use the FOR
READ ONLY clause for queries that are used in positioned UPDATE or DELETE
statements.
The OPTIMIZE FOR clause declares the intent to retrieve only a subset of the
result or to give priority to retrieving only the first few rows. The optimizer can
then prefer access plans that minimize the response time for retrieving the first few
rows. In addition, the number of rows that are sent to the client as a single block
are bounded by the value of “n” in the OPTIMIZE FOR clause. Thus the
OPTIMIZE FOR clause affects both how the server retrieves the qualifying rows
from the database by the server, and how it returns the qualifying rows to the
client.
For example, suppose you are querying the employee table for the employees with
the highest salary on a regular basis.
SELECT LASTNAME,FIRSTNAME,EMPNO,SALARY
FROM EMPLOYEE
ORDER BY SALARY DESC
You have defined a descending index on the SALARY column. However, since
employees are ordered by employee number, the salary index is likely to be very
poorly clustered. To avoid many random synchronous I/Os, the optimizer would
probably choose to use the list prefetch access method, which requires sorting the
row identifiers of all rows that qualify. This sort causes a delay before the first
qualifying rows can be returned to the application. To prevent this delay, add the
OPTIMIZE FOR clause to the statement as follows:
In this case, the optimizer probably chooses to use the SALARY index directly
because only the twenty employees with the highest salaries are retrieved.
Regardless of how many rows might be blocked, a block of rows is returned to the
client every twenty rows.
With the OPTIMIZE FOR clause the optimizer favors access plans that avoid bulk
operations or interrupt the flow of rows, such as sorts. You are most likely to
influence an access path by using OPTIMIZE FOR 1 ROW. Using this clause might
have the following effects:
v Join sequences with composite inner tables are less likely because they require a
temporary table.
v The join method might change. A nested loop join is the most likely choice,
because it has low overhead cost and is usually more efficient to retrieve a few
rows.
v An index that matches the ORDER BY clause is more likely because no sort is
required for the ORDER BY.
v List prefetch is less likely because this access method requires a sort.
v Sequential prefetch is less likely because of the understanding that only a small
number of rows is required.
v In a join query, the table with the columns in the ORDER BY clause is likely to
be picked as the outer table if an index on the outer table provides the ordering
needed for the ORDER BY clause.
Although the OPTIMIZE FOR clause applies to all optimization levels, it works
best for optimization class 3 and higher because classes below 3 use Greedy join
enumeration method. This method sometimes results in access plans for multi-table
joins that do not lend themselves to quick retrieval of the first few rows.
The OPTIMIZE FOR clause does not prevent you from retrieving all the qualifying
rows. If you do retrieve all qualifying rows, the total elapsed time might be
significantly greater than if the optimizer had optimized for the entire answer set.
If a packaged application uses the call level interface (DB2 CLI or ODBC), you can
use the OPTIMIZEFORNROWS keyword in the [Link] configuration file to
have DB2 CLI automatically append an OPTIMIZE FOR clause to the end of each
query statement.
When data is selected from nicknames, results may vary depending on data source
support. If the data source referenced by the nickname supports the OPTIMIZE
FOR clause and the DB2 optimizer pushes down the entire query to the data
source, then the clause is generated in the remote SQL sent to the data source. If
the data source does not support this clause or if the optimizer decides that the
least-cost plan is local execution, the OPTIMIZE FOR clause is applied locally. In
this case, the DB2 optimizer prefers access plans that minimize the response time
for retrieving the first few rows of a query, but the options available to the
optimizer for generating plans are slightly limited and performance gains from the
OPTIMIZE FOR clause may be negligible.
The FETCH FIRST n ROWS ONLY clause sets the maximum number of rows that
can be retrieved. Limiting the result table to the first several rows can improve
performance. Only n rows are retrieved regardless of the number of rows that the
result set might otherwise contain.
If you specify both the FETCH FIRST clause and the OPTIMIZE FOR clause, the
lower of the two values affects the communications buffer size. For optimization
purposes the two values are independent of each other.
When you declare a cursor with the DECLARE CURSOR statement that includes
the WITH HOLD clause, open cursors remain open when the transaction is
committed and all locks are released, except locks that protect the current cursor
position of open WITH HOLD cursors.
If the transaction is rolled back, all open cursors are closed and all locks are
released and LOB locators are freed.
Related concepts:
v “Query tuning guidelines” on page 81
v “Efficient SELECT statements” on page 83
Note: The block of rows that you specify is a number of pages in memory. It is not
a multi-dimensional (MDC) table block, which is physically mapped to an
extent on disk.
Row blocking levels are specified by the following arguments to the BIND or PREP
commands:
UNAMBIG
Blocking occurs for read-only cursors and cursors not specified as “FOR
UPDATE OF”. Ambiguous cursors are treated as updateable.
ALL Blocking occurs for read-only cursors and cursors not specified as “FOR
UPDATE OF”. Ambiguous cursors are treated as read-only.
NO Blocking does not occur for any cursors. Ambiguous cursors are treated as
read-only.
Prerequisites:
Procedure:
Note: If you use the FETCH FIRST n ROWS ONLY clause or the OPTIMIZE FOR n
ROWS clause in a SELECT statement, the number of rows per block will be
the minimum of the following:
v The value calculated in the above formula
v The value of n in the FETCH FIRST clause
v The value of n in the OPTIMIZE FOR clause
Related reference:
v “aslheapsz - Application support layer heap size” on page 358
v “rqrioblk - Client I/O block size” on page 360
Note: The optimization class that the optimizer uses might eliminate the need for
some fine tuning because the SQL compiler can rewrite the SQL code into
more efficient forms.
Note that the optimizer choice of an access plan is also affected by other factors,
including environmental considerations and system catalog statistics. If you
conduct benchmark testing of the performance of your applications, you can find
out what adjustments might improve the access plan.
Related tasks:
v “Specifying row blocking to reduce overhead” on page 80
The most common application of sampling is for aggregate queries such as AVG,
SUM, and COUNT, where reasonably accurate answers of the aggregates can be
obtained from a sample of the data. Sampling can also be used to obtain a random
subset of the actual rows in a table for auditing purposes or to speed up data
mining and analysis tasks.
System page-level sampling is similar to row-level sampling, except that pages are
sampled and not rows. A page is included in the sample with a probability of
P/100. If a page is included, all of the rows in that page are included.
The best sampling method for a particular task will be determined by a user’s time
constraints and the desired degree of accuracy.
To execute a query on a random sample of data from a table, you can use the
TABLESAMPLE clause of the table-reference clause in a SQL statement. To specify
the method of sampling, use the keywords BERNOULLI or SYSTEM.
Related reference:
v “Subselect” in the SQL Reference, Volume 1
Related concepts:
84 Administration Guide: Performance
v “Guidelines for restricting select statements” on page 77
v “Query tuning guidelines” on page 81
Compound SQL is supported in stored procedures, which are also known as DARI
routines, and in the following application development processes:
v Embedded static SQL
v DB2 Call Level Interface
v JDBC
Dynamic compound statements can also use several flow logic statements, such as
the FOR statement, the IF statement, the ITERATE statement, and the WHILE
statement.
Related concepts:
v “Query tuning guidelines” on page 81
Character-conversion guidelines
Data conversion might be required to map data between application and database
code pages when your application and database do not use the same code page.
Because mapping and data conversion require additional overhead application
performance improves if the application and database use the same code page or
the identity collating sequence.
The conversion function and conversion tables or DBCS conversion APIs that the
database manager uses when it converts multi-byte code pages depends on the
operating system environment.
Note: Character string conversions between multi-byte code pages, such as DBCS
with EUC, might increase or decrease length of a string. In addition, code
points assigned to different characters in the PC DBCS, EUC, and UCS-2
code sets might produce different results when same characters are sorted.
Many characters in both the Japanese and Traditional Chinese EUC code pages
require special methods of managing database and client application support for
graphic data, which require double byte characters. Graphic data from these EUC
code pages is stored and manipulated using the UCS-2 code set.
Related concepts:
v “Guidelines for analyzing where a federated query is evaluated” on page 182
Related reference:
v “Conversion tables for code pages 923 and 924” in the Administration Guide:
Planning
v “Conversion table files for euro-enabled code pages” in the Administration Guide:
Planning
Processing a single SQL statement for a remote database requires sending two
transmissions: one request and one receive. Because an application contains many
SQL statements it requires many transmissions to complete its work.
However, when a database client uses a stored procedure that encapsulates many
SQL statements, it requires only two transmissions for the entire process.
Stored procedures usually run in processes separate from the database agents. This
separation requires the stored procedure and agent processes to communicate
through a router. However, a special kind of stored procedure that runs in the
agent process might improve performance, although it carries significant risks of
corrupting data and databases.
These risky stored procedures are those created as not fenced. For a not-fenced
stored procedure, nothing separates the stored procedure from the database control
structures that the database agent uses. If a DBA wants to ensure that the stored
procedure operations will not accidentally or maliciously damage the database
control structures, the not fenced option is omitted.
Because of the risk of damaging your database, use not fenced stored procedures
only when you need the maximum possible performance benefits. In addition,
make absolutely sure that the procedure is well coded and has been thoroughly
tested before allowing it to run as a not-fenced stored procedure. If a fatal error
occurs while running a not-fenced stored procedure, the database manager
determines whether the error occurred in the application or database manager
code and performs the appropriate recovery.
A not-fenced stored procedure can corrupt the database manager beyond recovery,
possibly resulting in lost data and the possibility of a corrupt database. Exercise
Related concepts:
v “Query tuning guidelines” on page 81
If a query is compiled with DEGREE = ANY, the database manager chooses the
degree of intra-partition parallelism based on a number of factors including the
number of processors and the characteristics of the query. The actual degree used
at runtime may be lower than the number of processors depending on these
factors and the amount of activity on the system. Parallelism may be lowered
before query execution if the system is heavily utilized. This occurs because
intra-partition parallelism aggressively uses system resources to reduce the elapsed
time of the query, which may adversely affect the performance of other database
users.
Related concepts:
v “Explain tools” on page 190
v “Optimization strategies for intra-partition parallelism” on page 173
Related reference:
v “max_querydegree - Maximum query degree of parallelism” on page 450
v “intra_parallel - Enable intra-partition parallelism” on page 449
v “dft_degree - Default degree” on page 431
| The REOPT bind option specifies whether or not to have DB2® optimize an access
| path using values for host variables, parameter markers, and special registers.
| REOPT values are specified by the following arguments to the BIND, PREP and
| REBIND commands:
| REOPT NONE
| The access path for a given SQL statement containing host variables,
| parameter markers, or special registers will not be optimized using real
| values for these variables; The default estimates for the these variables are
| used instead. This plan is cached and will be used subsequently. This is the
| default behavior.
| REOPT ONCE
| The access path for a given SQL statement will be optimized using the real
| values of the host variables, parameter markers, or special registers when
| the query is first executed. This plan is cached and used subsequently.
| REOPT ALWAYS
| The access path for a given SQL statement will always be compiled and
| reoptimized using the values of the host variables, parameter markers, or
| special registers known at each execution time.
| Related concepts:
| v “Effects of REOPT on static SQL” in the Application Development Guide:
| Programming Client Applications
| v “Effects of REOPT on dynamic SQL” in the Application Development Guide:
| Programming Client Applications
For more information about factors that affect the SQL optimizer directly, see
Chapter 5, “System catalog statistics,” on page 95.
When you tune applications and make changes to the environment, rebind your
applications to ensure that the best access plan is used.
In a partitioned database, depending on the size of the table, spreading data over
more partitions reduces the estimated time (or cost) to execute a query. The
number of tables, the size of the tables, the location of the data in those tables, and
the type of query, such as whether a join is required, all affect the cost of the query.
Related concepts:
v “Join strategies in partitioned databases” on page 164
v “Join methods in partitioned databases” on page 165
v “Partitions in a partitioned database” on page 282
where:
- 0.5 represents an average overhead of one half rotation
- Rotational latency is calculated in milliseconds for each full rotation, as
follows:
(1 / RPM) * 60 * 1000
where you:
v Divide by rotations per minute to get minutes per rotation
v Multiply by 60 seconds per minute
v Multiply by 1000 milliseconds per second.
As an example, let the rotations per minute for the disk be 7 200. Using the
rotational-latency formula,this would produce:
(1 / 7200) * 60 * 1000 = 8.328 milliseconds
which can then be used in the calculation of the OVERHEAD estimate with
an assumed average seek time of 11 milliseconds:
OVERHEAD = 11 + (0.5 * 8.328)
= 15.164
where:
- spec_rate represents the disk specification for the transfer rate, in MB per
second
- Divide by spec_rate to get seconds per MB
- Multiply by 1000 milliseconds per second
- Divide by 1 024 000 bytes per MB
- Multiply by the page size in bytes (for example, 4 096 bytes for a 4 KB
page)
As an example, suppose the specification rate for the disk is 3 MB per second.
This would produce the following calculation
TRANSFERRATE = (1 / 3) * 1000 / 1024000 * 4096
= 1.333248
The following shows an example of the syntax to change the characteristics of the
RESOURCE table space:
ALTER TABLESPACE RESOURCE
PREFETCHSIZE 64
OVERHEAD 19.3
TRANSFERRATE 0.9
Related concepts:
v “Catalog statistics tables” on page 106
v “The SQL compiler process” on page 133
v “Illustration of the DMS table-space address map” on page 17
Note: You must install the distributed join installation option and set the database
manager parameter federated to YES before you can create servers and
specify server options.
The server option values that you specify affect query pushdown analysis, global
optimization and other aspects of federated database operations. For example, in
the CREATE SERVER statement, you can specify performance statistics as server
option values, such as the cpu_ratio option, which specifies the relative speeds of
the CPUs at the data source and the federated server. You might also set the
io_ratio option to a value that indicates the relative rates of the data I/O divides at
the source and the federated server. When you execute the CREATE SERVER
statement, this data is added to the catalog view [Link], and
the optimizer uses it in developing its access plan for the data source. If a statistic
changes (as might happen, for instance, if the data source CPU is upgraded), use
the ALTER SERVER statement to update [Link] with this
change. The optimizer then uses the new information the next time it chooses an
access plan for the data source.
Related reference:
v “ALTER SERVER statement” in the SQL Reference, Volume 2
v “[Link] catalog view” in the SQL Reference, Volume 1
v “federated - Federated database system support” on page 458
v “CREATE SEQUENCE statement” in the SQL Reference, Volume 2
Catalog statistics
When the SQL compiler optimizes SQL query plans, its decisions are heavily
influenced by statistical information about the size of the database tables and
indexes. The optimizer also uses information about the distribution of data in
specific columns of tables and indexes if these columns are used to select rows or
join tables. The optimizer uses this information to estimate the costs of alternative
access plans for each query.
In addition to table size and data distribution information, you can also collect
statistical information about the cluster ratio of indexes, the number of leaf pages
in indexes, the number of table rows that overflow their original pages, and the
number of filled and empty pages in a table. You use this information to decide
when to reorganize tables and indexes.
Statistical information is collected for specific tables and indexes in the local
database when you execute the RUNSTATS utility. The collected statistics are
stored in the system catalog tables.
Statistics are collected only for the table partition that resides on the partition
where you execute the utility or the first partition in the database partition group
that contains the table.
Note: Because the RUNSTATS utility does not support use of nicknames, you
update statistics differently for federated database queries. If queries access
a federated database, execute RUNSTATS for the tables in all databases, then
drop and recreate the nicknames that access remote tables to make the new
statistics available to the optimizer.
Consider these tips to improve the efficiency of RUNSTATS and the usefulness of
the collected statistics:
v Collect statistics only for the columns used to join tables or in the WHERE,
GROUP BY, and similar clauses of queries. If these columns are indexed, you
can specify the columns with the ONLY ON KEY COLUMNS clause for the
RUNSTATS command.
v Customize the values for num_freqvalues and num_quantiles for specific tables and
specific columns in tables.
v Collect DETAILED index statistics with the SAMPLE DETAILED clause to
reduce the amount of background calculation performed for detailed index
statistics. The SAMPLE DETAILED clause reduces the time required to collect
statistics, and produces adequate precision in most cases.
Note: You can perform a RUNSTATS on a declared temporary table, but the
resulting statistics are not stored in the system catalogs because declared
temporary tables do not have catalog entries. However, the statistics are
stored in memory structures that represent the catalog information for
declared temporary tables. In some cases, therefore, it might be useful to
perform a RUNSTATS on these tables.
Related concepts:
v “Catalog statistics tables” on page 106
v “Statistical information that is collected” on page 111
v “Catalog statistics for modeling and what-if planning” on page 124
v “Statistics for modeling production databases” on page 125
v “General rules for updating catalog statistics manually” on page 127
Related tasks:
v “Collecting catalog statistics” on page 98
Related reference:
v “num_freqvalues - Number of frequent values retained” on page 434
v “num_quantiles - Number of quantiles for columns” on page 435
v “RUNSTATS Command” in the Command Reference
Note: RUNSTATS only collects statistics for tables on the partition from which you
execute it. The RUNSTATS results from this partition are extrapolated to the
other partitions. If the database partition from which you execute
RUNSTATS does not contain a table partition, the request is sent to the first
database partition in the database partition group that holds a partition for
the table.
To improve RUNSTATS performance and save disk space used to store statistics,
consider specifying only the columns for which data distribution statistics should
be collected.
Ideally, you should rebind application programs after running statistics. The query
optimizer might choose a different access plan if it has new statistics.
If you do not have enough time to collect all of the statistics at one time, you
might run RUNSTATS to update statistics on only a few tables and indexes at a
time, rotating through the set of tables. If inconsistencies occur as a result of
activity on the table between the periods where you run RUNSTATS with a
selective partial update, then a warning message (SQL0437W, reason code 6) is
issued during query optimization. For example, you first use RUNSTATS to gather
table distribution statistics. Subsequently, you use RUNSTATS to gather index
statistics. If inconsistencies occur as a result of activity on the table and are
detected during query optimization, the warning message is issued. When this
happens, you should run RUNSTATS again to update distribution statistics.
To ensure that the index statistics are synchronized with the table, execute
RUNSTATS to collect both table and index statistics at the same time. Index
statistics retain most of the table and column statistics collected from the last run
of RUNSTATS. If the table has been modified extensively since the last time its
table statistics were gathered, gathering only the index statistics for that table will
leave the two sets of statistics out of synchronization on all nodes.
Related concepts:
v “Catalog statistics” on page 95
v “Optimizer use of distribution statistics” on page 115
v “Automatic statistics collection” on page 104
Related tasks:
v “Collecting catalog statistics” on page 98
v “Collecting distribution statistics for specific columns” on page 99
v “Collecting index statistics” on page 100
Prerequisites:
You must connect to the database that contains the tables and indexes and have
one of the following authorization levels:
v sysadm
v sysctrl
v sysmaint
v dbadm
v CONTROL privilege on the table
Procedure:
Note: RUNSTATS only collects statistics for tables on the partition from which
you execute it. The RUNSTATS results from this partition are
extrapolated to the other partitions. If the database partition from which
you execute RUNSTATS does not contain a table partition, the request is
sent to the first database partition in the database partition group that
holds a partition for the table.
3. When RUNSTATS is complete, issue a COMMIT statement to release locks.
4. Rebind packages that access tables and indexes for which you have regenerated
statistical information.
To use a graphical user interface to specify options and collect statistics, use the
Control Center.
Related tasks:
v “Collecting distribution statistics for specific columns” on page 99
v “Collecting index statistics” on page 100
v “Determining when to reorganize tables” on page 240
Related reference:
v “RUNSTATS Command” in the Command Reference
In the following steps, the database is assumed to be sales and to contain the table
customers, with indexes custidx1 and custidx2.
Prerequisites:
You must connect to the database that contains the tables and indexes and have
one of the following authorization levels:
v sysadm
v sysctrl
v sysmaint
v dbadm
v CONTROL privilege on the table
Note: RUNSTATS only collects statistics for tables on the partition from which you
execute it. The RUNSTATS results from this partition are extrapolated to the
other partitions. If the database partition from which you execute
RUNSTATS does not contain a table partition, the request is sent to the first
database partition in the database partition group that holds a partition for
the table.
Procedure:
You can also use the Control Center to collect distribution statistics.
Related concepts:
v “Catalog statistics tables” on page 106
Related tasks:
v “Collecting catalog statistics” on page 98
v “Collecting index statistics” on page 100
In the following steps, the database is assumed be sales and to contain the table
customers, with indexes custidx1 and custidx2.
Prerequisites:
You must connect to the database that contains the tables and indexes and have
one of the following authorization levels:
v sysadm
v sysctrl
v sysmaint
v dbadm
v CONTROL privilege on the table
Executing RUNSTATS with the SAMPLED DETAILED option requires 2MB of the
statistics heap. Allocate an additional 488 4K pages to the stat_heap_sz database
configuration parameter setting for this additional memory requirement. If the
heap appears to be too small, RUNSTATS returns an error before it attempts to
collect statistics.
Procedure:
You can also use the Control Center to collect index and table statistics.
Related concepts:
v “Catalog statistics” on page 95
v “Catalog statistics tables” on page 106
v “Statistical information that is collected” on page 111
v “Detailed index statistics” on page 120
Related tasks:
v “Collecting catalog statistics” on page 98
| Starting in Version 8.2, the RUNSTATS command provides the option to collect
| statistics on a sample of the data in the table by using the TABLESAMPLE option.
| This feature can increase the efficiency of statistics collection since sampling uses
| only a subset of the data. At the same time, the sampling methods ensure a high
| level of accuracy.
| There are two ways to specify how the sample is to be collected. The BERNOULLI
| method samples the data at the level of the row. During a full table scan of the
| data pages each row is considered in turn and is selected based on probability P as
| specified by the numeric parameter. It is only on these selected rows that statistics
| will be collected. In a similar manner, the SYSTEM method samples the data at the
| page-level. Thus, each page is selected on probability P and rejected with
| probability 1-P/100.
| Related concepts:
| v “Data sampling in SQL queries” on page 82
| Related reference:
| v “RUNSTATS Command” in the Command Reference
| This feature simplifies statistics collection by allowing you to store the options that
| you specify when you issue the RUNSTATS command so that you can collect the
| same statistics repeatedly on a table without having to re-type the command
| options.
| You can register or update a statistics profile with or without actually collecting
| statistics. For example, to register a profile and collect statistics at the same time,
| issue the RUNSTATS command with the SET PROFILE option. To register a profile
| only, without actually collecting statistics, issue the RUNSTATS command with the
| SET PROFILE ONLY option.
| To collect statistics using a statistics profile that you have already registered, issue
| the RUNSTATS command, specifying only the name of the table and the USE
| PROFILE option.
| To see what options are currently specified in the statistics profile for a particular
| table, you can query the catalog tables with the following select statement, where
| tablename is the name of the table that you want the profile for:
| SELECT STATISTICS_PROFILE FROM [Link] WHERE NAME = tablename
| This feature can be used with the automatic statistics collection feature, which
| automatically schedules statistics maintenance based on the information contained
| within the automatically generated statistics profile.
| To enable this feature, you need to have already enabled automatic table
| maintenance by setting the appropriate configuration parameters . The
| AUTO_STATS_PROF configuration parameter activates the collection of query
| feedback data, and the AUTO_PROF_UPD configuration parameter activates the
| generation of a statistics profile for use by automatic statistics collection.
| Note: Automatic statistics profile generation can only be activated in DB2 serial
| mode, and is blocked for queries in federated, SMP, or MPP environments.
| Note: There is some performance overhead associated with monitoring the queries
| and storing the query feedback data in the feedback warehouse.
| where:
| toolname
| Specifies the name of the tool whose objects are to be created or dropped.
| In this case ″ASP″ or ″AUTO STATS PROFILING″.
| action Specifies the action to be taken: ’C’ for create, ’D’ for drop.
| tablespacename
| The name of the table space in which the the feedback warehouse tables
| will be created. This input parameter is optional. If it is not specified, the
| default user space will be used.
| schemaname
| The name of the schema with which the objects will be created or dropped.
| This parameter is currently not used.
| For example, to create the feedback warehouse in table space ″A″ enter: call
| SYSINSTALLOBJECTS ("ASP", ’C’, "A", "")
| Related concepts:
| v “Automatic statistics collection” on page 104
| Related tasks:
| v “Using automatic statistics collection” on page 105
| Related concepts:
| v “Collecting statistics using a statistics profile” on page 102
| Related tasks:
| v “Using automatic statistics collection” on page 105
| Related reference:
| v “util_impact_lim - Instance impact policy” on page 463
| v “autonomic_switches - Automatic maintenance switches” on page 437
| Procedure:
| You can turn this feature on using either the graphical user interface tools or the
| command line interface.
| v To set up your database for automatic statistics collection using the graphical
| user interface tools:
| 1. Open the Configure Automatic Maintenance wizard either from the Control
| Center by right-clicking on a database object or from the Health Center by
| right-clicking on the database instance that you want to configure for
| automatic statistics collection. Select Configure Automatic Maintenance from
| the pop-up window.
| 2. Within this wizard, you can enable automatic statistics collection, specify the
| tables that you want to automatically collect statistics from, and specify a
| maintenance window for the execution of the RUNSTATS utility.
| 3. OPTIONAL: To enable the automatic statistics profile generation, set the
| following two configuration parameters to ″ON″ using the command line
| interface:
| Related concepts:
| v “Collecting statistics using a statistics profile” on page 102
| v “Automatic statistics collection” on page 104
| Related reference:
| v “autonomic_switches - Automatic maintenance switches” on page 437
Statistics collected
This section lists the catalog statistics tables and describes the use of the fields in
these tables. Sections that follow the statistics tables descriptions explain the kind
of data that can be collected and stored in the tables.
Note: The multicolumn distribution statistics listed in the following two tables are
not collected by RUNSTATS. You can update them manually, however.
Table 21. Multicolumn Distribution Statistics ([Link] and [Link])
Statistic Description RUNSTATS Option
Table Indexes
TYPE F = frequency value Yes No
Q = quantile value
ORDINAL Ordinal number of the Yes No
column in the group
SEQNO Sequence number n that Yes No
represents the nth TYPE
value
COLVALUE the data value as a Yes No
character literal or a null
value
Q = quantile value
SEQNO Sequence number n that Yes No
represents the nth TYPE
value
VALCOUNT If TYPE = F, VALCOUNT is Yes No
the number of occurrences
of COLVALUEs for the
column group identified by
this SEQNO.
If TYPE = Q, VALCOUNT
is the number of rows
whose value is less than or
equal to COLVALUEs for
the column group with this
SEQNO.
DISTCOUNT If TYPE = Q, this column Yes No
contains the number of
distinct values that are less
than or equal to
COLVALUEs for the
column group with this
SEQNO. Null if
unavailable.
Related concepts:
v “Catalog statistics” on page 95
v “Statistical information that is collected” on page 111
v “Statistics for user-defined functions” on page 123
v “Statistics for modeling production databases” on page 125
For each column in the table and the first column in the index key:
v The cardinality of the column
v The average length of the column
v The second highest value in the columns
v The second lowest value in the column
v The number of NULLs in the column
If you request detailed statistics for an index, you also store finer information
about the degree of clustering of the table to the index and the page fetch
estimates for different buffer sizes.
You can also collect the following kinds statistics about tables and indexes:
v Data distribution statistics
The optimizer uses data distribution statistics to estimate efficient access plans
for tables in which data is not evenly distributed and columns have a significant
number of duplicate values.
v Detailed index statistics
The optimizer uses detailed index statistics to determine how efficient it is to
access a table through an index.
v Sub-element statistics
The optimizer uses sub-element statistics for LIKE predicates, especially those
that search for a pattern embedded within a string, such as LIKE %disk%.
Related concepts:
v “Catalog statistics” on page 95
v “Catalog statistics tables” on page 106
Related tasks:
v “Collecting catalog statistics” on page 98
v “Collecting distribution statistics for specific columns” on page 99
v “Collecting index statistics” on page 100
Distribution statistics
You can collect two kinds of data distribution statistics:
v Frequency statistics
These statistics provide information about the column and the data value with
the highest number of duplicates, the next highest number of duplicate values,
and so on to the level specified by the value of the num_freqvalues database
configuration parameter. To disable collection of frequent-value statistics, set
num_freqvalues to 0.
Note: If you specify larger num_freqvalues and num_quantiles values, more CPU
resources and memory, as specified by the stat_heap_sz database
configuration parameter, are required when you execute RUNSTATS.
To decide whether distribution statistics should be created and updated for a given
table, consider the following two factors:
v Whether applications use static or dynamic SQL.
Distribution statistics are most useful for dynamic SQL and static SQL that does
not use host variables. When using SQL with host variables, the optimizer
makes limited use of distribution statistics.
v Whether data in columns is distributed uniformly.
Create distribution statistics if at least one column in the table has a highly
“non-uniform” distribution of data and the column appears frequently in
equality or range predicates; that is, in clauses such as the following:
WHERE C1 = KEY;
WHERE C1 IN (KEY1, KEY2, KEY3);
WHERE (C1 = KEY1) OR (C1 = KEY2) OR (C1 = KEY3);
WHERE C1 <= KEY;
WHERE C1 BETWEEN KEY1 AND KEY2;
C1
0.0
5.1
6.3
7.1
8.2
8.4
8.5
9.1
93.6
100.0
To help the optimizer deal with duplicate values, create both quantile and
frequent-value statistics.
You might collect statistics based only on index data in the following situations:
v A new index has been created since the RUNSTATS utility was run and you do
not want to collect statistics again on the table data.
v There have been many changes to the data that affect the first column of an
index.
To determine the precision with which distribution statistics are stored, you specify
the database configuration parameters, num_quantiles and num_freqvalues. You can
also specify these parameters as RUNSTATS options when you collect statistics for
a table or for columns. The higher you set these values, the greater precision
RUNSTATS uses when it create and updates distribution statistics. However,
greater precision requires greater use of resources, both during RUNSTATS
execution and in the storage required in the catalog tables.
For most databases, specify between 10 and 100 for the num_freqvalues database
configuration parameter. Ideally, frequent-value statistics should be created such
that the frequencies of the remaining values are either approximately equal to each
other or negligible compared to the frequencies of the most frequent values. The
database manager might collect less than this number, because these statistics will
only be collected for data values that occur more than once. If you need to collect
only quantile statistics, set num_freqvalues to zero.
To set the number of quantiles, specify between 20 and 50 as the setting of the
num_quantiles database configuration parameter. A rough rule of thumb for
determining the number of quantiles is:
v Determine the maximum error that is tolerable in estimating the number of rows
of any range query, as a percentage, P
Related concepts:
v “Catalog statistics” on page 95
v “Catalog statistics tables” on page 106
v “Optimizer use of distribution statistics” on page 115
v “Extended examples of distribution-statistics use” on page 116
Related tasks:
v “Collecting catalog statistics” on page 98
v “Collecting distribution statistics for specific columns” on page 99
Related reference:
v “num_freqvalues - Number of frequent values retained” on page 434
v “num_quantiles - Number of quantiles for columns” on page 435
If you do not execute RUNSTATS with the WITH DISTRIBUTION clause, the
catalog statistics tables contain information only about the size of the table and the
highest and lowest values in the table, the degree of clustering of the table to any
of its indexes, and the number of distinct values in indexed columns.
Unless it has additional information about the distribution of values between the
low and high values, the optimizer assumes that data values are evenly
distributed. If data values differ widely from each other, are clustered in some
parts of the range, or contain many duplicate values, the optimizer will choose a
less than optimal access plan.
The optimizer needs to estimate the number of rows containing a column value
that satisfies an equality or range predicate in order to select the least expensive
access plan. The more accurate the estimate, the greater the likelihood that the
optimizer will choose the optimal access plan. For example, consider the query
SELECT C1, C2
FROM TABLE1
WHERE C1 = ’NEW YORK’
AND C2 <= 10
Assume that there is an index on both C1 and C2. One possible access plan is to
use the index on C1 to retrieve all rows with C1 = ’NEW YORK’ and then check each
When distribution statistics are not available but RUNSTATS has been executed
against a table, the only information available to the optimizer is the
second-highest data value (HIGH2KEY), second-lowest data value (LOW2KEY),
number of distinct values (COLCARD), and number of rows (CARD) for a column.
The number of rows that satisfy an equality or range predicate is then estimated
under the assumption that the frequencies of the data values in a column are all
equal and the data values are evenly spread out over the interval (LOW2KEY,
HIGH2KEY). Specifically, the number of rows satisfying an equality predicate C1 =
KEY is estimated as CARD/COLCARD, and the number of rows satisfying a range
predicate C1 BETWEEN KEY1 AND KEY2 is estimated as:
KEY2 - KEY1
------------------- x CARD (1)
HIGH2KEY - LOW2KEY
These estimates are accurate only when the true distribution of data values in a
column is reasonably uniform. When distribution statistics are unavailable and
either the frequencies of the data values differ widely from each other or the data
values are clustered in a few sub-intervals of the interval (LOW_KEY,HIGH_KEY),
the estimates can be off by orders of magnitude and the optimizer may choose a
less than optimal access plan.
When distribution statistics are available, the errors described above can be greatly
reduced by using frequent-value statistics to compute the number of rows that
satisfy an equality predicate and using frequent-value statistics and quantiles to
compute the number of rows that satisfy a range predicate.
Related concepts:
v “Catalog statistics” on page 95
v “Distribution statistics” on page 112
v “Extended examples of distribution-statistics use” on page 116
Related tasks:
v “Collecting distribution statistics for specific columns” on page 99
If frequent-value statistics are available, the optimizer can use these statistics to
choose an appropriate access plan, as follows:
v If KEY is one of the N most frequent values, then the optimizer uses the
frequency of KEY that is stored in the catalog.
where CARD is the number of rows in the table, COLCARD is the cardinality of
the column and NUM_FREQ_ROWS is the total number of rows with a value
equal to one of the N most frequent values.
For example, consider a column (C1) for which the frequency of the data values is
as follows:
If frequent-value statistics based on only the most frequent value (that is, N = 1)
are available, for this column, the number of rows in the table is 50 and the
column cardinality is 5. For the predicate C1 = 3, exactly 40 rows satisfy it. If the
optimizer assumes that data is evenly distributed, it estimates the number of rows
that satisfy the predicate as 50/5 = 10, with an error of -75%. If the optimizer can
use frequent-value statistics, the number of rows is estimated as 40, with no error.
Using the frequent value statistics (N = 1), the optimizer will estimate the number
of rows containing this value using the formula (2) given above, for example:
(50 - 40)
--------- = 3
(5 - 1)
The following explanations of quantile statistics use the term “K-quantile”. The
K-quantile for a column is the smallest data value, V, such that at least “K” rows
have data values less than or equal to V. To computer a K-quantile, sort the rows in
the column according to increasing data values; the K-quantile is the data value in
the Kth row of the sorted column.
C
0.0
5.1
6.3
7.1
8.2
8.4
8.5
9.1
93.6
100.0
and suppose that K-quantiles are available for K = 1, 4, 7, and 10, as follows:
K K-quantile
1 0.0
4 7.1
7 8.5
10 100.0
First consider the predicate C <= 8.5. For the data given above, exactly 7 rows
satisfy this predicate. Assuming a uniform data distribution and using formula (1)
from above, with KEY1 replaced by LOW2KEY, the number of rows that satisfy the
predicate is estimated as:
8.5 - 5.1
---------- x 10 *= 0
93.6 - 5.1
If quantile statistics are available, the optimizer estimates the number of rows that
satisfy this same predicate (C <= 8.5) by locating 8.5 as the highest value in one of
the quantiles and estimating the number of rows by using the corresponding value
of K, which is 7. In this case, the error is reduced to 0.
Now consider the predicate C <= 10. Exactly 8 rows satisfy this predicate. If the
optimizer must assume a uniform data distribution and use formula (1), the
number of rows that satisfy the predicate is estimated as 1, an error of -87.5%.
Unlike the previous example, the value 10 is not one of the stored K-quantiles.
However, the optimizer can use quantiles to estimate the number of rows that
satisfy the predicate as r_1 + r_2, where r_1 is the number of rows satisfying the
predicate C <= 8.5 and r_2 is the number of rows satisfying the predicate C > 8.5
AND C <= 10. As in the above example, r_1 = 7. To estimate r_2 the optimizer uses
linear interpolation:
10 - 8.5
r_2 *= ---------- x (number of rows with value > 8.5 and <= 100.0)
100 - 8.5
10 - 8.5
r_2 *= ---------- x (10 - 7)
100 - 8.5
The final estimate is r_1 + r_2 *= 7, and the error is only -12.5%.
Quantiles improves the accuracy of the estimates in the above examples because
the real data values are ″clustered″ in the range 5 - 10, but the standard estimation
formulas assume that the data values are spread out evenly between 0 and 100.
The use of quantiles also improves accuracy when there are significant differences
in the frequencies of different data values. Consider a column having data values
with the following frequencies:
Suppose that K-quantiles are available for K = 5, 25, 75, 95, and 100:
K K-quantile
5 20
25 40
75 50
95 70
100 80
Also suppose that frequent value statistics are available based on the 3 most
frequent values.
Consider the predicate C BETWEEN 20 AND 30. From the distribution of the data
values, you can see that exactly 10 rows satisfy this predicate. Assuming a uniform
data distribution and using formula (1), the number of rows that satisfy the
predicate is estimated as:
30 - 20
------- x 100 = 25
70 - 30
Using frequent-value statistics and quantiles, the number of rows that satisfy the
predicate is estimated as r_1 + r_2, where r_1 is the number of rows that satisfy
the predicate (C = 20) and r_2 is the number of rows that satisfy the predicate C >
20 AND C <= 30. Using formula (2), r_1 is estimated as:
100 - 80
-------- = 5
7 - 3
Related concepts:
v “Catalog statistics” on page 95
v “Distribution statistics” on page 112
v “Optimizer use of distribution statistics” on page 115
v “Rules for updating distribution statistics manually” on page 129
Related tasks:
v “Collecting distribution statistics for specific columns” on page 99
Note: When you collect detailed index statistics, RUNSTATS takes longer and
requires more memory and CPU processing. The SAMPLED DETAILED
option, for which information calculated only for a statistically significant
number of entries, requires 2MB of the statistics heap. Allocate an additional
488 4K pages to the stat_heap_sz database configuration parameter setting for
this additional memory requirement. If the heap appears to be too small,
RUNSTATS returns an error before attempting to collect statistics.
The DETAILED statistics provide concise information about the number of physical
I/Os required to access the data pages of a table if a complete index scan is
performed under different buffer sizes. As RUNSTATS scans the pages of the
index, it models the different buffer sizes, and gathers estimates of how often a
page fault occurs. For example, if only one buffer page is available, each new page
referenced by the index results in a page fault. In a worse case, each row might
reference a different page, resulting in at most the same number of I/Os as rows in
the indexed table. At the other extreme, when the buffer is big enough to hold the
entire table (subject to the maximum buffer size), then all table pages are read
once. As a result, the number of physical I/Os is a monotone, non-increasing
function of the buffer size.
You should collect DETAILED index statistics when queries reference columns that
are not included in the index. In addition, DETAILED index statistics should be
used in the following circumstances:
v The table has multiple unclustered indexes with varying degrees of clustering
v The degree of clustering is non-uniform among the key values
v The values in the index are updated non-uniformly
Related concepts:
v “Catalog statistics” on page 95
v “Catalog statistics tables” on page 106
Related tasks:
v “Collecting catalog statistics” on page 98
v “Collecting index statistics” on page 100
Sub-element statistics
If tables contain columns that contain sub-fields or sub-elements separated by
blanks, and queries reference these columns in WHERE clauses, you should collect
sub-element statistics to ensure the best access plans.
For queries that specify LIKE predicates on such columns using the % match_all
character:
SELECT .... FROM DOCUMENTS WHERE KEYWORDS LIKE ’%simulation%’
it is often beneficial for the optimizer to know some basic statistics about the
sub-element structure of the column.
The following statistics are collected when you execute RUNSTATS with the LIKE
STATISTICS clause:
where xxxxxx is any string of characters; that is, any LIKE predicate whose search
value starts with a % character. (It might or might not end with a % character).
These are referred to as ″wildcard LIKE predicates″. For all predicates, the
optimizer has to estimate how many rows match the predicate. For wildcard LIKE
predicates, the optimizer assumes that the COLUMN being matched contains a
series of elements concatenated together, and it estimates the length of each
element based on the length of the string, excluding leading and trailing %
characters.
Note: RUNSTATS might take longer if you use the LIKE STATISTICS clause. For
example, RUNSTATS might take between 15% and 40%, and longer on a
table with five character columns, if the DETAILED and DISTRIBUTION
options are not used. If either the DETAILED or the DISTRIBUTION option
is specified, the overhead percentage is less, even though the absolute
amount of overhead is the same. If you are considering using this option,
you should assess this overhead against improvements in query
performance.
Related concepts:
v “Catalog statistics” on page 95
v “Catalog statistics tables” on page 106
Related tasks:
v “Collecting catalog statistics” on page 98
v “Collecting distribution statistics for specific columns” on page 99
The following table provides information about the statistic columns for which you
can provide estimates to improve performance:
Table 25. Function Statistics ([Link] and [Link])
Statistic Description
IOS_PER_INVOC Estimated number of read/write requests
executed each time a function is executed.
INSTS_PER_INVOC Estimated number of machine instructions
executed each time a function is executed.
IOS_PER_ARGBYTE Estimated number of read/write requests
executed per input argument byte.
INSTS_PER_ARGBYTES Estimated number of machine instructions
executed per input argument byte.
PERCENT_ARGBYTES Estimated average percent of input
argument bytes that the function will
actually process.
INITIAL_IOS Estimated number of read/write requests
executed only the first/last time the function
is invoked.
INITIAL_INSTS Estimated number of machine instructions
executed only the first/last time the function
is invoked.
CARDINALITY Estimated number of rows generated by a
table function.
For example, consider a UDF (EU_SHOE) that converts an American shoe size to
the equivalent European shoe size. (These two shoe sizes could be UDTs.) For this
UDF, you might set the statistic columns as follows:
v INSTS_PER_INVOC: set to the estimated number of machine instructions
required to:
– Invoke EU_SHOE
– Initialize the output string
– Return the result.
v INSTS_PER_ARGBYTE: set to the estimated number of machine instructions
required to convert the input string into a European shoe size.
v PERCENT_ARGBYTES: set to 100 indicating that the entire input string is to be
converted
v INITIAL_INSTS, IOS_PER_INVOC, IOS_PER_ARGBYTE, and INITIAL_IOS: set
each to 0, since this UDF only performs computations.
Related concepts:
v “Catalog statistics tables” on page 106
v “General rules for updating catalog statistics manually” on page 127
Do not manually update statistics on a production system. If you do, the optimizer
might not choose the best access plan for production queries that contain dynamic
SQL.
Requirements
You must have explicit DBADM authority for the database to modify statistics for
tables and indexes and their components. That is, your user ID is recorded as
having DBADM authority in the [Link] table. Belonging to a DBADM
group does not explicitly provide this authority. A DBADM can see statistics rows
for all users, and can execute SQL UPDATE statements against the views defined
in the SYSSTAT schema to update the values of these statistical columns.
A user without DBADM authority can see only those rows which contain statistics
for objects over which they have CONTROL privilege. If you do not have DBADM
authority, you can change statistics for individual database objects if you have the
following privileges for each object:
The following shows an example of updating the table statistics for the
EMPLOYEE table:
UPDATE [Link]
SET CARD = 10000,
NPAGES = 1000,
FPAGES = 1000,
OVERFLOW = 2
WHERE TABSCHEMA = ’userid’
AND TABNAME = ’EMPLOYEE’
You must be careful when manually updating catalog statistics. Arbitrary changes
can seriously alter the performance of subsequent queries. Even in a
non-production database that you are using for testing or modeling, you can use
any of the following methods to refresh updates you applied to these tables and
bring the statistics to a consistent state:
v ROLLBACK the unit of work in which the changes have been made (assuming
the unit of work has not been committed).
v Use the RUNSTATS utility to recalculate and refresh the catalog statistics.
v Update the catalog statistics to indicate that statistics have not been gathered.
(For example, setting column NPAGES to -1 indicates that the number-of-pages
statistic has not been collected.)
v Replace the catalog statistics with the data they contained before you made any
changes. This method is possible only if you used the db2look tool to capture the
statistics before you made any changes.
In some cases, the optimizer may determine that some particular statistical value
or combination of values is not valid. It will use default values and issue a
warning. Such circumstances are rare, however, since most of the validation is
done when updating the statistics.
Related concepts:
v “Catalog statistics” on page 95
v “Catalog statistics tables” on page 106
v “Statistics for user-defined functions” on page 123
v “Statistics for modeling production databases” on page 125
A productivity tool, db2look, can be run against the production database to generate
the update statements required to make the catalog statistics of the test database
You can recreate database data objects, including tables, views, indexes, and other
objects in a database, by extracting DDL statements with db2look -e. You can run
the command processor script created from this command against another database
to recreate the database. You can use -e option and the -m option together in a
script that re-creates the database and sets the statistics.
After running the update statements produced by db2look against the test system,
the test system can be used to validate the access plans to be generated in
production. Since the optimizer uses the type and configuration of the table spaces
to estimate I/O costs, the test system must have the same table space geometry or
layout. That is, the same number of containers of the same type, either SMS or
DMS.
For more information on how to use this productivity tool, type the following on a
command line:
db2look -h
The Control Center also provides an interface to the db2look utility called “Generate
SQL - Object Name”. Using the Control Center allows the results file from the
utility to be integrated into the Script Center. You can also schedule the db2look
command from the Control Center. One difference when using the Control Center
is that only single table analysis can be done as opposed to a maximum of thirty
tables in a single call using the db2look command. You should also be aware that
LaTex and Graphical outputs are not supported from the Control Center.
You can also run the db2look utility against an OS/390 or z/OS database. The
db2look utility extracts the DDL and UPDATE statistics statements for OS/390
objects. This is very useful if you would like to extract OS/390 or z/OS objects and
re-create them in a DB2® Universal Database (UDB) database. /p>
There are some differences between the DB2 UDB statistics and the OS/390
statistics. The db2look utility performs the appropriate conversions from DB2 for
OS/390 or z/OS to DB2 UDB when this is applicable and sets to a default value
(-1) the DB2 UDB statistics for which a DB2 for OS/390 counterpart does not exist.
Here is how the db2look utility maps the DB2 for OS/390 or z/OS statistics to DB2
UDB statistics. In the discussion below, “UDB_x” stands for a DB2 UDB statistics
column; and, “S390_x” stands for a DB2 for OS/390 or z/OS statistics column.
1. Table Level Statistics.
UDB_CARD = S390_CARDF
UDB_NPAGES = S390_NPAGES
There is no S390_FPAGES. However, DB2 for OS/390 or z/OS has another
statistics called PCTPAGES which represents the percentage of active table
space pages that contain rows of the table. So it is possible to calculate
UDB_FPAGES based on S390_NPAGES and S390_PCTPAGES as follows:
UDB_FPAGES=(S390_NPAGES * 100)/S390_PCTPAGES
UDB_COLCARD = S390_COLCARDF
UDB_HIGH2KEY = S390_HIGH2KEY
UDB_LOW2KEY = S390_LOW2KEY
There is no S390_AVGCOLLEN to map to UDB_AVGCOLLEN so the db2look
utility just sets this to the default value:
UDB_AVGCOLLEN=-1
3. Index Level Statistics.
UDB_NLEAF = S390_NLEAF
UDB_NLEVELS = S390_NLEVELS
UDB_FIRSTKEYCARD= S390_FIRSTKEYCARD
UDB_FULLKEYCARD = S390_FULLKEYCARD
UDB_CLUSTERRATIO= S390_CLUSTERRATIO
The other statistics for which there are no OS/390 or z/OS counterparts are just
set to the default. That is:
UDB_FIRST2KEYCARD = -1
UDB_FIRST3KEYCARD = -1
UDB_FIRST4KEYCARD = -1
UDB_CLUSTERFACTOR = -1
UDB_SEQUENTIAL_PAGES = -1
UDB_DENSITY = -1
4. Column Distribution Statistics.
There are two types of statistics in DB2 for OS/390 or z/OS
[Link]. Type “F” for frequent values and type “C” for
cardinality. Only entries of type “F” are applicable to DB2 for UDB and these
are the ones that will be considered.
UDB_COLVALUE = S390_COLVALUE
UDB_VALCOUNT = S390_FrequencyF * S390_CARD
In addition, there is no column SEQNO in DB2 for OS/390
[Link]. Because this required for DB2 for UDB, db2look
generates one automatically.
Related concepts:
v “Catalog statistics” on page 95
v “Catalog statistics tables” on page 106
v “Catalog statistics for modeling and what-if planning” on page 124
v “General rules for updating catalog statistics manually” on page 127
Related reference:
v “db2look - DB2 Statistics and DDL Extraction Tool Command” in the Command
Reference
The most common checks you should make, before updating a catalog statistic, are:
1. Numeric™ statistics must be -1 or greater than or equal to zero.
2. Numeric statistics representing percentages (for example, CLUSTERRATIO in
[Link]) must be between 0 and 100.
Note: For row types, the table level statistics NPAGES, FPAGES, and OVERFLOW
are not updateable for a sub-table.
Related concepts:
v “Catalog statistics tables” on page 106
v “Statistics for user-defined functions” on page 123
v “Catalog statistics for modeling and what-if planning” on page 124
v “Statistics for modeling production databases” on page 125
v “Rules for updating column statistics manually” on page 128
v “Rules for updating distribution statistics manually” on page 129
v “Rules for updating table and nickname statistics manually” on page 130
v “Rules for updating index statistics manually” on page 131
– HIGH2KEY should be greater than LOW2KEY whenever there are more than
three distinct values in the corresponding column.
v The cardinality of a column (COLCARD statistic in [Link]) cannot
be greater than the cardinality of its corresponding table (CARD statistic in
[Link]).
v The number of nulls in a column (NUMNULLS statistic in [Link])
cannot be greater than the cardinality of its corresponding table (CARD statistic
in [Link]).
v No statistics are supported for columns with data types: LONG VARCHAR,
LONG VARGRAPHIC, BLOB, CLOB, DBCLOB.
Related concepts:
v “Catalog statistics” on page 95
v “Catalog statistics tables” on page 106
v “Catalog statistics for modeling and what-if planning” on page 124
v “General rules for updating catalog statistics manually” on page 127
Make sure that all the statistics in the catalog are consistent. Specifically, for each
column, the catalog entries for the frequent data statistics and quantiles must
satisfy the following constraints:
v Frequent value statistics (in the [Link] catalog). These constraints
include:
– The values in column VALCOUNT must be unchanging or decreasing for
increasing values of SEQNO.
– The number of values in column COLVALUE must be less than or equal to
the number of distinct values in the column, which is stored in column
COLCARD in catalog view [Link].
– The sum of the values in column VALCOUNT must be less than or equal to
the number of rows in the column, which is stored in column CARD in
catalog view [Link].
– In most cases, the values in the column COLVALUE should lie between the
second-highest and second-lowest data values for the column, which are
stored in columns HIGH2KEY and LOW2KEY, respectively, in catalog view
[Link]. There may be one frequent value greater than
HIGH2KEY and one frequent value less than LOW2KEY.
v Quantiles (in the [Link] catalog). These constraints include:
Suppose that distribution statistics are available for a column C1 with “R” rows
and you wish to modify the statistics to correspond to a column with the same
relative proportions of data values, but with “(F x R)” rows. To scale up the
frequent-value statistics by a factor of F, each entry in column VALCOUNT must
be multiplied by F. Similarly, to scale up the quantiles by a factor of F, each entry
in column VALCOUNT must be multiplied by F. If you do not follow these rules,
the optimizer might use the wrong filter factor and cause unpredictable
performance when you run the query.
Related concepts:
v “Catalog statistics” on page 95
v “Catalog statistics tables” on page 106
v “Catalog statistics for modeling and what-if planning” on page 124
v “General rules for updating catalog statistics manually” on page 127
When working within a federated database system, use caution when manually
providing or updating statistics on a nickname over a remote view. The statistical
information, such as the number of rows this nickname will return, might not
reflect the real cost to evaluate this remote view and thus might mislead the DB2®
optimizer. Situations that can benefit from statistics updates include remote views
defined on a single base table with no column functions applied on the SELECT
list. Complex views may require a complex tuning process which might require
that each query be tuned. Consider creating local views over nicknames instead so
the DB2 optimizer knows how to derive the cost of the view more accurately.
Related concepts:
where
NPAGES = 300
CARD = 10000
CLUSTERRATIO = -1
CLUSTERFACTOR = 0.9
2. CLUSTERRATIO and CLUSTERFACTOR (in [Link]) must adhere
to the following rules:
v Valid values for CLUSTERRATIO are -1 or between 0 and 100.
v Valid values for CLUSTERFACTOR are -1 or between 0 and 1.
v At least one of the CLUSTERRATIO and CLUSTERFACTOR values must be
-1 at all times.
v If CLUSTERFACTOR is a positive value, it must be accompanied by a valid
PAGE_FETCH_PAIR statistic.
3. The following rules apply to FIRSTKEYCARD, FIRST2KEYCARD,
FIRST3KEYCARD, FIRST4KEYCARD, and FULLKEYCARD:
v FIRSTKEYCARD must be equal to FULLKEYCARD for a single-column
index.
v FIRSTKEYCARD must be equal to COLCARD (in [Link]) for
the corresponding column.
Related concepts:
v “Catalog statistics” on page 95
v “Catalog statistics tables” on page 106
v “Catalog statistics for modeling and what-if planning” on page 124
v “General rules for updating catalog statistics manually” on page 127
The topics in this chapter provide more information about how the SQL compiler
compiles and optimizes SQL statements.
Parse Query
Check
Semantics
Rewrite
Query
Query
Graph
Pushdown Model
Analysis
Optimize
Access Plan
Generate
Executable Code
Execute Plan
Explain Executable
Tables Plan
The query graph model is an internal, in-memory database that represents the query
as it is processed in the steps described below:
1. Parse Query
The SQL compiler analyzes the SQL query to validate the syntax. If any syntax
errors are detected, the SQL compiler stops processing and returns the
appropriate SQL error to the application that submitted the query. When
parsing is complete, an internal representation of the query is created and
stored in the query graph model.
2. Check Semantics
The compiler ensures that there are no inconsistencies among parts of the
statement. As a simple example of semantic checking, the compiler verifies that
the data type of the column specified for the YEAR scalar function is a
datetime data type.
Note: Execute RUNSTATS at appropriate intervals on tables that change often. The
optimizer needs up-to-date statistical information about the tables and their
data to create the most efficient access plans. Rebind your application to
take advantage of updated statistics. If RUNSTATS is not executed or the
optimizer suspects that RUNSTATS was executed on empty or nearly empty
tables, it may either use defaults or attempt to derive certain statistics based
on the number of file pages used to store the table on disk (FPAGES). The
total number of occupied blocks is stored in the ACTIVE_BLOCKS column.
Related concepts:
v “Query rewriting methods and examples” on page 139
v “Data-access methods” on page 148
v “Predicate terminology” on page 154
v “Joins” on page 156
v “Effects of sorting and grouping” on page 171
v “Optimization strategies for intra-partition parallelism” on page 173
v “Materialized query tables” on page 176
v “Guidelines for analyzing where a federated query is evaluated” on page 182
v “Advantages of Deferred Binding” in the Application Development Guide:
Programming Client Applications
v “Optimization strategies for MDC tables” on page 175
The following configuration parameters or factors affect the access plan chosen by
the SQL compiler:
v The size of the buffer pools that you specified when you created or altered them.
When the optimizer chooses the access plan, it considers the I/O cost of fetching
pages from disk to the buffer pool and estimates the number of I/Os required to
satisfy a query. The estimate includes a prediction of buffer-pool usage, because
additional physical I/Os are not required to read rows in a page that is already
in the buffer pool.
The optimizer considers the value of the npages column in the BUFFERPOOLS
system catalog tables and, on partitioned databases, the
BUFFERPOOLDBPARTITION system catalog tables.
The I/O costs of reading the tables can have an impact on:
– How two tables are joined
– Whether an unclustered index will be used to read the data
v Default Degree (dft_degree)
The dft_degree configuration parameter specifies parallelism by providing a
default value for the CURRENT DEGREE special register and the DEGREE bind
option. A value of one (1) means no intra-partition parallelism. A value of minus
one (-1) means the optimizer determines the degree of intra-partition parallelism
based on the number of processors and the type of query.
v Default Query Optimization Class (dft_queryopt)
Although you can specify a query optimization class when you compile SQL
queries, you might set a default optimization degree.
Note: Intra-parallel processing does not occur unless you enable it by setting the
intra_parallel database configuration parameter.
v Average Number of Active Applications (avg_appls)
The SQL optimizer uses the avg_appls parameter to help estimate how much of
the buffer pool might be available at run-time for the access plan chosen. Higher
values for this parameter can influence the optimizer to choose access plans that
are more conservative in buffer pool usage. If you specify a value of 1, the
optimizer considers that the entire buffer pool will be available to the
application.
v Sort Heap Size (sortheap)
If the rows to be sorted occupy more than the space available in the sort heap,
several sort passes are performed, where each pass sorts a subset of the entire
set of rows. Each sort pass is stored in a temporary table in the buffer pool,
Related concepts:
v “The SQL compiler process” on page 133
Related reference:
v “max_querydegree - Maximum query degree of parallelism” on page 450
v “comm_bandwidth - Communications bandwidth” on page 456
v “sortheap - Sort heap size” on page 355
v “locklist - Maximum storage for lock list” on page 340
v “maxlocks - Maximum percent of lock list before escalation” on page 369
v “stmtheap - Statement heap size” on page 357
v “cpuspeed - CPU speed” on page 457
v “avg_appls - Average number of active applications” on page 378
v “dft_degree - Default degree” on page 431
Query rewriting
This section the ways in which the optimizer can rewrite queries to improve
performance.
To influence the number of query rewrite rules that are applied to an SQL
statement, change the optimization class. To see some of the results of the query
rewrite, use the Explain facility or Visual Explain.
Related concepts:
v “Compiler rewrite example: view merges” on page 140
v “Compiler rewrite example: DISTINCT elimination” on page 143
v “Compiler rewrite example: implied predicates” on page 144
v “Column correlation for multiple predicates” on page 145
During query rewrite, these two views could be merged to create the following
query:
SELECT [Link], [Link], [Link], [Link], [Link]
FROM EMPLOYEE E1,
EMPLOYEE E2
WHERE [Link] = [Link]
AND [Link] > 17
AND [Link] > 35000
By merging the SELECT statements from the two views with the user-written
SELECT statement, the optimizer can consider more choices when selecting an
access plan. In addition, if the two views that have been merged use the same base
table, additional rewriting may be performed.
The SQL compiler will take a query containing a subquery, such as:
SELECT EMPNO, FIRSTNME, LASTNAME, PHONENO
FROM EMPLOYEE
WHERE WORKDEPT IN
(SELECT DEPTNO
FROM DEPARTMENT
WHERE DEPTNAME = ’OPERATIONS’)
In this query, the SQL compiler can eliminate the join and simplify the query to:
SELECT EMPNO, FIRSTNME, LASTNAME, EDLEVEL, SALARY
FROM EMPLOYEE
WHERE EDLEVEL > 17
AND SALARY > 35000
becomes
SELECT LASTNAME, SALARY
FROM EMPLOYEE
WHERE WORKDEPT NOT NULL
Note that in this situation, even if users know that the query can be re-written,
they may not be able to do so because they do not have access to the underlying
tables. They may only have access to the view shown above. Therefore, this type of
optimization has to be performed within the database manager.
Using multiple functions within a query can generate several calculations which
take time. Reducing the number of calculations to be done within the query results
in an improved plan. The SQL compiler takes a query using multiple functions
such as:
SELECT SUM(SALARY+BONUS+COMM) AS OSUM,
AVG(SALARY+BONUS+COMM) AS OAVG,
COUNT(*) AS OCOUNT
FROM EMPLOYEE;
This rewrite reduces the query from 2 sums and 2 counts to 1 sum and 1 count.
Related concepts:
v “The SQL compiler process” on page 133
v “Query rewriting methods and examples” on page 139
In the above example, since the primary key is being selected, the SQL compiler
knows that each row returned will already be unique. In this case, the DISTINCT
key word is redundant. If the query is not rewritten, the optimizer would need to
build a plan with the necessary processing, such as a sort, to ensure that the
columns are distinct.
Altering the level at which a predicate is normally applied can result in improved
performance. For example, given the following view which provides a list of all
employees in department “D11”:
CREATE VIEW D11_EMPLOYEE
(EMPNO, FIRSTNME, LASTNAME, PHONENO, SALARY, BONUS, COMM)
AS SELECT EMPNO, FIRSTNME, LASTNAME, PHONENO, SALARY, BONUS, COMM
FROM EMPLOYEE
WHERE WORKDEPT = ’D11’
The query rewrite stage of the compiler will push the predicate LASTNAME =
’BROWN’ down into the view D11_EMPLOYEE. This allows the predicate to be
applied sooner and potentially more efficiently. The actual query that could be
executed in this example is:
SELECT FIRSTNME, PHONENO
FROM EMPLOYEE
WHERE LASTNAME = ’BROWN’
AND WORKDEPT = ’D11’
Example - Decorrelation
In a partitioned database environment, the SQL compiler can rewrite the following
query:
Find all the employees who are working on programming projects and are
underpaid.
SELECT [Link], [Link], [Link], [Link],
[Link]+[Link]+[Link] AS COMPENSATION
FROM EMPLOYEE E, PROJECT P
WHERE [Link] = [Link]
Since this query is correlated, and since both PROJECT and EMPLOYEE are
unlikely to be partitioned on PROJNO, the broadcast of each project to each
database partition is possible. In addition, the subquery would have to be
evaluated many times.
The rewritten SQL query computes the AVG_COMP per project (AVG_PRE_PROJ) and
can then broadcast the result to all database partitions containing the EMPLOYEE
table.
Related concepts:
v “The SQL compiler process” on page 133
v “Query rewriting methods and examples” on page 139
As a result of this rewrite, the optimizer can consider additional joins when it is
trying to select the best access plan for the query.
In addition to the above predicate transitive closure, query rewrite also derives
additional local predicates based on the transitivity implied by equality predicates.
For example, the following query lists the names of the departments whose
department number is greater than “E00” and the employees who work in those
departments.
SELECT EMPNO, LASTNAME, FIRSTNAME, DEPTNO, DEPTNAME
FROM EMPLOYEE EMP,
DEPARTMENT DEPT
WHERE [Link] = [Link]
AND [Link] > ’E00’
For this query, the rewrite stage adds the following implied predicate:
[Link] > ’E00’
As a result of this rewrite, the optimizer reduces the number of rows to be joined.
Example - OR to IN Transformations
Note: In some cases, the database manager might convert an IN predicate to a set
of OR clauses so that index ORing might be performed.
Related concepts:
v “The SQL compiler process” on page 133
v “Query rewriting methods and examples” on page 139
For example, consider a manufacturer who makes products from raw material of
various colors, elasticities and qualities. The finished product has the same color
and elasticity as the raw material from which it is made. The manufacturer issues
the query:
This query returns the names and raw material quality of all products. There are
two join predicates:
[Link] = [Link]
[Link] = [Link]
When the optimizer chooses a plan for executing this query, it calculates how
selective each of the two predicates is. It assumes that they are independent, which
means that all variations of elasticity occur for each color, and that conversely for
each level of elasticity there is raw material of every color. It then estimate the
overall selectivity of the pair of predicates by using catalog statistic information for
each table on the number of levels of elasticity and the number of different colors.
Based on this estimate, it may choose, for example, a nested loop join in preference
to a merge join, or vice versa.
However, it may be that these two predicates are not independent. For example, it
may be that the highly elastic materials are available in only a few colors, and the
very inelastic materials are only available in a few other colors that are different
from the elastic ones. Then the combined selectivity of the predicates eliminates
fewer rows so the query will return more rows. Consider the extreme case, in
which there is just one level of elasticity for each color and vice versa. Now either
one of the predicates logically could be omitted entirely since it is implied by the
other. The optimizer might no longer choose the best plan. For example, it might
choose a nested loop join plan when the merge join would be faster.
With other database products, database administrators have tried to solve this
performance problem by updating statistics in the catalog to try to make one of the
predicates appear to be less selective, but this approach can cause unwanted side
effects on other queries.
| The DB2® UDB optimizer attempts to detect and compensate for correlation of join
| predicates if you define an index on those columns or if you collect and maintain
| group column statistics on the appropriate columns.
For example, in elasticity example above, you might define a unique index
covering either:
[Link], [Link]
or
[Link], [Link]
or both.
For the optimizer to detect correlation, the non-include columns of this index must
be only the correlated columns. The index may also contain include columns to
allow index-only scans. If there are more than two correlated columns in join
predicates, make sure that you define the unique index to cover all of them. In
many cases, the correlated columns in one table are its primary key. Because a
primary key is always unique, you do not need to define another unique index.
After creating appropriate indexes, ensure that statistics on tables are up to date
and that they have not been manually altered from the true values for any reason,
such as to attempt to influence the optimizer.
| Column group statistics are collected using the ″ON COLUMNS″ option of
| RUNSTATS. For example, to collect the column group statistics on
| [Link] and [Link], issue the following RUNSTATS
| command:
| RUNSTATS ON TABLE product ON COLUMNS ((color, elasticity))
| If an index exists on the two columns MAKE and MODEL or column group
| statistics are gathered, the optimizer uses the statistical information about the index
| or columns to determine the combined number of distinct values and adjust the
| selectivity or cardinality estimation for correlation between the two columns. If
| such predicates are not join predicates, the optimizer does not need a unique index
| to make the adjustment.
Related concepts:
v “The SQL compiler process” on page 133
v “Query rewriting methods and examples” on page 139
| Related concepts:
Data-access methods
When it compiles an SQL statement, the SQL optimizer estimates the execution
cost of different ways of satisfying the query. Based on its estimates, the optimizer
selects an optimal access plan. An access plan specifies the order of operations
required to resolve an SQL statement. When an application program is bound, a
package is created. This package contains access plans for all of the static SQL
statements in that application program. Access plans for dynamic SQL statements
are created at the time that the application is executed.
To produce the results that the query requests, rows are selected depending on the
terms of the predicate, which are usually stated in a WHERE clause. The selected
rows in accessed tables are joined to produce the result set, and the result set
might be further processed by grouping or sorting the output.
Related concepts:
v “The SQL compiler process” on page 133
v “Data access through index scans” on page 148
v “Types of index access” on page 151
v “Index access and cluster ratios” on page 153
If indexes are created with the ALLOW REVERSE SCANS option, scans may also
be performed in the direction opposite to that with which they were defined.
Note: The optimizer chooses a table scan if no appropriate index has been created
or if an index scan would be more costly. An index scan might be more
To determine whether an index can be used for a particular query, the optimizer
evaluates each column of the index starting with the first column to see if it can be
used to satisfy equality and other predicates in the WHERE clause. A predicate is an
element of a search condition in a WHERE clause that expresses or implies a
comparison operation. Predicates that can be used to delimit the range of an index
scan in the following cases:
v Tests for equality against a constant, a host variable, an expression that evaluates
to a constant, or a keyword
v Tests for “IS NULL” or “IS NOT NULL”
v Tests for equality against a basic subquery, which is a subquery that does not
contain ANY, ALL, or SOME, and the subquery does not have a correlated
column reference to its immediate parent query block (that is, the SELECT for
which this subquery is a subselect).
v Tests for strict and inclusive inequality.
The following examples illustrate when an index might be used to limit a range:
v Consider an index with the following definition:
INDEX IX1: NAME ASC,
DEPT ASC,
MGR DESC,
SALARY DESC,
YEARS ASC
In this case, the following predicates might be used to limit the range of the scan
of index IX1:
WHERE NAME = :hv1
AND DEPT = :hv2
or
WHERE MGR = :hv1
AND NAME = :hv2
AND DEPT = :hv3
Note that in the second WHERE clause, the predicates do not have to be
specified in the same order as the key columns appear in the index. Although
the examples use host variables, other variables such as parameter markers,
expressions, or constants would have the same effect.
v Consider a single index created using the ALLOW REVERSE SCANS parameter.
Such indexes support scans in the direction defined when the index was created
as well as in the opposite or reverse direction. The statement might look
something like this:
CREATE INDEX iname ON tname (cname DESC) ALLOW REVERSE SCANS
In this case, the index (iname) is formed based on DESCending values in cname.
By allowing reverse scans, although the index on the column is defined for scans
in descending order, a scan can be done in ascending order. The actual use of
the index in both directions is not controlled by you but by the optimizer when
creating and considering access plans.
This is because there is a key column (MGR) separating these columns from the
first two index key columns, so the ordering would be off. However, once the
range is determined by the NAME = :hv1 and DEPT = :hv2 predicates, the
remaining predicates can be evaluated against the remaining index key columns.
Certain inequality predicates can delimit the range of an index scan. There are two
types of inequality predicates:
v Strict inequality predicates
The strict inequality operators used for range delimiting predicates are greater
than ( > ) and less than ( < ).
Only one column with strict inequality predicates is considered for delimiting a
range for an index scan. In the following example, the predicates on the NAME
and DEPT columns can be used to delimit the range, but the predicate on the
MGR column cannot be used.
WHERE NAME = :hv1
AND DEPT > :hv2
AND DEPT < :hv3
AND MGR < :hv4
v Inclusive inequality predicates
The following are inclusive inequality operators that can be used for range
delimiting predicates:
– >= and <=
– BETWEEN
– LIKE
For delimiting a range for an index scan, multiple columns with inclusive
inequality predicates will be considered. In the following example, all of the
predicates can be used to delimit the range of the index scan:
WHERE NAME = :hv1
AND DEPT >= :hv2
AND DEPT <= :hv3
AND MGR <= :hv4
To further illustrate this example, suppose that :hv2 = 404, :hv3 = 406, and :hv4
= 12345. The database manager will scan the index for all of departments 404
and 405, but it will stop scanning department 406 when it reaches the first
manager that has an employee number (MGR column) greater than 12345.
If the query requires output in sorted order, an index might be used to order the
data if the ordering columns appear consecutively in the index, starting from the
first index key column. Ordering or sorting can result from operations such as
ORDER BY, DISTINCT, GROUP BY, “= ANY” subquery, “> ALL” subquery, “<
ALL” subquery, INTERSECT or EXCEPT, UNION. An exception to this is when the
For this query, the index might be used to order the rows because NAME and
DEPT will always be the same values and will thus be ordered. That is, the
preceding WHERE and ORDER BY clauses are equivalent to:
WHERE NAME = ’JONES’
AND DEPT = ’D93’
ORDER BY NAME, DEPT, MGR
A unique index can also be used to truncate a sort-order requirement. Consider the
following index definition and ORDER BY clause:
UNIQUE INDEX IX0: PROJNO ASC
SELECT PROJNO, PROJNAME, DEPTNO
FROM PROJECT
ORDER BY PROJNO, PROJNAME
Additional ordering on the PROJNAME column is not required because the IX0
index ensures that PROJNO is unique. This uniqueness ensures that there is only
one PROJNAME value for each PROJNO value.
Related concepts:
v “Data-access methods” on page 148
v “Index structure” on page 23
v “Types of index access” on page 151
v “Index access and cluster ratios” on page 153
Index-Only Access
In some cases, all of the required data can be retrieved from the index without
accessing the table. This is known as an index-only access.
The following query can be satisfied by accessing only the index, and without
reading the base table:
Often, however, required columns that do not appear in the index. To obtain the
data for these columns, the table rows must be read. To allow the optimizer to
choose an index-only access, create a unique index with include columns. For
example, consider the following index definition:
CREATE UNIQUE INDEX IX1 ON EMPLOYEE
(NAME ASC)
INCLUDE (DEPT, MGR, SALARY, YEARS)
This index enforces uniqueness of the NAME column and also stores and
maintains data for DEPT, MGR, SALARY, and YEARS columns, which allows the
following query to be satisfied by accessing only the index:
SELECT NAME, DEPT, MGR, SALARY
FROM EMPLOYEE
WHERE NAME=’SMITH’
The optimizer can choose to scan multiple indexes on the same table to satisfy the
predicates of a WHERE clause. For example, consider the following two index
definitions:
INDEX IX2: DEPT ASC
INDEX IX3: JOB ASC,
YEARS ASC
Scanning index IX2 produces a list of row IDs (RIDs) that satisfy the DEPT = :hv1
predicate. Scanning index IX3 produces a list of RIDs satisfying the JOB = :hv2 AND
YEARS >= :hv3 predicate. These two lists of RIDs are combined and duplicates
removed before the table is accessed. This is known as index ORing.
Index ORing may also be used for predicates specified in the IN clause, as in the
following example:
WHERE DEPT IN (:hv1, :hv2, :hv3)
Although the purpose of index ORing is to eliminate duplicate RIDs, the objective
of index ANDing is to find common RIDs. Index ANDing might occur with
applications that create multiple indexes on corresponding columns in the same
table and a query using multiple AND predicates is run against that table. Multiple
index scans against each indexed column in such a query produce values which
are hashed to create bitmaps. The second bitmap is used to probe the first bitmap
to generate the qualifying rows that are fetched to create the final returned data
set.
In this example, scanning index IX4 produces a bitmap satisfying the SALARY
BETWEEN 20000 AND 30000 predicate. Scanning IX5 and probing the bitmap for IX4
results in the list of qualifying RIDs that satisfy both predicates. This is known as
“dynamic bitmap ANDing”. It occurs only if the table has sufficient cardinality and
the columns have sufficient values in the qualifying range, or sufficient duplication
if equality predicates are used.
Additional sort heap space is required when dynamic bitmaps are used in access
plans. When sheapthres is set to be relatively close to sortheap (that is, less than a
factor of two or three times per concurrent query), dynamic bitmaps with multiple
index access must work with much less memory than the optimizer anticipated.
The solution is to increase the value of sheapthres relative to sortheap.
Note: The optimizer does not combine index ANDing and index ORing in
accessing a single table.
Unlike standard tables, a range clustered table does not require a physical index
that maps a key value to a row like a traditional B-tree index. Instead, it leverages
the sequential nature of the column domain and uses a functional mapping to
generate the location of a given row in a table. In the simplest example of this
mapping, the first key value in the range is the first row in the table, and the
second value in the range is the second row in the table, and so on.
The optimizer uses the range-clustered property of the table to generate access
plans based on a perfectly clustered index whose only cost is computing the range
clustering function. The clustering of rows within the table is guaranteed because
range clustered tables retain their original key value ordering.
Related concepts:
v “Advantages and disadvantages of indexes” on page 244
v “Data-access methods” on page 148
v “Data access through index scans” on page 148
v “Index access and cluster ratios” on page 153
If index clustering statistics are not available, the optimizer uses default values,
which assume poor clustering of the data to the index.
The degree to which the data is clustered with respect to the index can have a
significant impact on performance and you should try to keep one of the indexes
on the table close to 100 percent clustered.
In general, only one index can be one hundred percent clustered, except in those
cases where the keys are a superset of the keys of the clustering index or where
there is de facto correlation between the key columns of the two indexes.
When you reorganize an table, you can specify an index that will be used to
cluster the rows and attempt to preserve this characteristic during insert
processing. Because updates and inserts may make the table less well clustered in
relation to the index, you might need to periodically reorganize the table. To
reduce the frequency of reorganization on a table that has frequent changes due to
INSERTs, UPDATEs, and DELETES, use the PCTFREE parameter when you alter a
table. This allows for additional inserts to be clustered with the existing data.
Related concepts:
v “Index performance tips” on page 248
v “Types of index access” on page 151
Predicate terminology
A user application requests a set of rows from the database with an SQL statement
that specifies qualifiers for the specific rows to be returned as the result set. These
qualifiers usually appear in the WHERE clause of the query. Such qualifiers are
called predicates. Predicates can be grouped into four categories that are determined
by how and when the predicate is used in the evaluation process. The categories
are listed below, ordered in terms of performance from best to worst:
1. Range delimiting predicates
2. Index SARGable predicates
3. Data SARGable predicates
4. Residual predicates.
Range delimiting predicates limit the scope of an index scan. They provide start
and stop key values for the index search. Index SARGable predicates cannot limit
the scope of a search, but can be evaluated from the index because the columns
involved in the predicate are part of the index key. For example, consider the
following index:
INDEX IX1: NAME ASC,
DEPT ASC,
MGR DESC,
SALARY DESC,
YEARS ASC
The first two predicates (NAME = :hv1, DEPT = :hv2) are range-delimiting
predicates, while YEARS > :hv5 is an index SARGable predicate.
The optimizer uses the index data when it evaluates these predicates instead of
reading the base table. These index SARGable predicates reduce the set of rows that
need to be read from the table, but they do not affect the number of index pages
that are accessed.
Predicates that cannot be evaluated by the index manager, but can be evaluated by
data management services are called data SARGable predicates. These predicates
usually require accessing individual rows from a table. If required, Data
Management Services retrieve the columns needed to evaluate the predicate, as
well as any others to satisfy the columns in the SELECT list that could not be
obtained from the index.
Residual Predicates
Residual predicates require more I/O costs than accessing a table. They might have
the following characteristics:
v Use correlated subqueries
v Use quantified subqueries, which contain ANY, ALL, SOME, or IN clauses
v Read LONG VARCHAR or LOB data, which is stored in a file that is separate
from the table
Such predicates are evaluated by Relational Data Services.
Sometimes predicates that are applied only to the index must be reapplied when
the data page is accessed. For example, access plans that use index ORing or index
ANDing always reapply the predicates as residual predicates when the data page
is accessed.
Related concepts:
v “The SQL compiler process” on page 133
Joins
A join is the process of combining information from two or more tables based on
some common domain of information. Rows from one table are paired with rows
from another table when information in the corresponding rows match on the
joining criterion.
Table1 Table2
PROJ PROJ_ID PROJ_ID NAME
A 1 1 Sam
B 2 3 Joe
C 3 4 Mary
D 4 1 Sue
2 Mike
To join Table1 and Table2 where the ID columns have the same values, use the
following SQL statement:
SELECT PROJ, x.PROJ_ID, NAME
FROM TABLE1 x, TABLE2 y
WHERE x.PROJ_ID = y.PROJ_ID
When two tables are joined, one table is selected as the outer table and the other as
the inner. The outer table is accessed first and is scanned only once. Whether the
inner table is scanned multiple times depends on the type of join and the indexes
that are present. Even if a query joins more than two tables, the optimizer joins
only two tables at a time. If necessary, temporary tables are created to hold
intermediate results.
You can provide explicit join operators, such as INNER or LEFT OUTER JOIN to
determine how tables are used in the join. Before you alter a query in this way,
however, you should allow the optimizer to determine how to join the tables. Then
analyze query performance to decide whether to add join operators.
Related concepts:
v “Join methods” on page 157
v “Join strategies in partitioned databases” on page 164
v “Join methods in partitioned databases” on page 165
v “Join information” on page 566
Join methods
The optimizer can choose one of three basic join strategies when queries require
tables to be joined.
v Nested-loop join
v Merge join
v Hash join
Nested-Loop Join
When it evaluates a nested loop join, the optimizer also decides whether to sort the
outer table before performing the join. If it orders the outer table, based on the join
columns, the number of read operations to access pages from disk for the inner
table might be reduced, because they are more likely to be be in the buffer pool
already. If the join uses a highly clustered index to access the inner table and if the
outer table has been sorted, the number of index pages accessed might be
minimized.
In addition, if the optimizer expects that the join will make a later sort more
expensive, it might also choose to perform the sort before the join. A later sort
might be required to support a GROUP BY, DISTINCT, ORDER BY or merge join.
Merge Join
Merge join, sometimes known as merge scan join or sort merge join, requires a
predicate of the form [Link] = [Link]. This is called an equality join
predicate. Merge join requires ordered input on the joining columns, either through
index access or by sorting. A merge join cannot be used if the join column is a
LONG field column or a large object (LOB) column.
To perform a merge join, the database manager performs the following steps:
1. Read the first row from T1. The value for A is “2”.
2. Scan T2 until a match is found, and then join the two rows.
3. Keep scanning T2 while the columns match, joining rows.
4. When the “3” in T2 is read, go back to T1 and read the next row.
5. The next value in T1 is “3”, which matches T2, so join the rows.
6. Keep scanning T2 while the columns match, joining rows.
7. The end of T2 is reached.
8. Go back to T1 to get the next row — note that the next value in T1 is the same
as the previous value from T1, so T2 is scanned again starting at the first “3” in
T2. The database manager remembers this position.
Hash Join
First, the designated INNER table is scanned and the rows copied into memory
buffers drawn from the sort heap specified by the sortheap database configuration
parameter. The memory buffers are divided into partitions based on a hash value
that is computed on the columns of the join predicates. If the size of the INNER
table exceeds the available sort heap space, buffers from selected partitions are
written to temporary tables.
When the inner table has been processed, the second, or OUTER, table is scanned
and its rows are matched to rows from the INNER table by first comparing the
hash value computed for the columns of the join predicates. If the hash value for
the OUTER row column matches the hash value of the INNER row column, the
actual join predicate column values are compared.
OUTER table rows that correspond to partitions not written to a temporary table
are matched immediately with INNER table rows in memory. If the corresponding
INNER table partition was written to a temporary table, the OUTER row is also
written to a temporary table. Finally, matching pairs of partitions from temporary
tables are read, and the hash values of their rows are matched, and the join
predicates are checked.
For the full performance benefits of hash joins, you might need to change the value
of the sortheap database configuration parameter and the sheapthres database
manager configuration parameter.
Hash-join performance is best if you can avoid hash loops and overflow to disk. To
tune hash-join performance, estimate the maximum amount of memory available
for sheapthres, then tune the sortheap parameter. Increase its setting until you avoid
as many hash loops and disk overflows as possible, but do not reach the limit
specified by the sheapthres parameter.
Increasing the sortheap value should also improve performance of queries that have
multiple sorts.
Related concepts:
v “Joins” on page 156
v “Join strategies in partitioned databases” on page 164
v “Join methods in partitioned databases” on page 165
v “Join information” on page 566
Related reference:
v “sortheap - Sort heap size” on page 355
v “sheapthres - Sort heap threshold” on page 354
Star-Schema Joins
The only exceptions occur when the optimization class is set to 9 or in the special
case of star schemas. A star schema contains a central table called the fact table and
the other tables are called dimension tables. The dimension tables all have only a
single join that attaches them to the fact table, regardless of the query. Each
dimension table contains additional values that expand information about a
particular column in the fact table. A typical query consists of multiple local
predicates that reference values in the dimension tables and contains join
predicates connecting the dimension tables to the fact table. For these queries it
might be beneficial to compute the Cartesian product of multiple small dimension
tables before accessing the large fact table. This technique is beneficial when
multiple join predicates match a multi-column index.
DB2 can recognize queries against databases designed with star schemas that have
at least two dimension tables and can increase the search space to include possible
plans that compute the Cartesian product of dimension tables. If the plan that
computes the Cartesian products has the lowest estimated cost, it is selected by the
optimizer.
The star schema join strategy discussed above assumes that primary key indexes
are used in the join. Another scenario involves foreign key indexes. If the foreign
key columns in the fact table are single-column indexes and there is a relatively
high selectivity across all dimension tables, the following star join technique can be
used:
1. Process each dimension table by:
v Performing a semi-join between the dimension table and the foreign key
index on the fact table
v Hashing the row ID (RID) values to dynamically create a bitmap.
2. Use AND predicates against the previous bitmap for each bitmap.
3. Determine the surviving RIDs after processing the last bitmap.
4. Optionally sort these RIDs.
5. Fetch a base table row.
6. Rejoin the fact table with each of its dimension tables, accessing the columns in
dimension tables that are needed for the SELECT clause.
7. Reapply the residual predicates.
The dynamic bitmaps created and used by star join techniques require sort heap
memory, the size of which is specified by the Sort Heap Size (sortheap) database
configuration parameter.
Composite Tables
Related concepts:
v “Joins” on page 156
v “Join methods” on page 157
v “Join strategies in partitioned databases” on page 164
v “Join methods in partitioned databases” on page 165
To update the content of the replicated materialized query table, run the following
statement:
REFRESH TABLE R_EMPLOYEE;
The following example calculates sales by employee, the total for the department,
and the grand total:
SELECT [Link], [Link], SUM([Link])
FROM department AS d, employee AS e, sales AS s
WHERE s.sales_person = [Link]
AND [Link] = [Link]
GROUP BY ROLLUP([Link], [Link])
ORDER BY [Link], [Link];
Instead of using the EMPLOYEE table, which is on only one database partition, the
database manager uses the R_EMPLOYEE table, which is replicated on each of the
database partitions where the SALES tables is stored. The performance
enhancement occurs because the employee information does not have to be moved
across the network to each database partition to calculate the join.
Replicated materialized query tables can also assist in the collocation of joins. For
example, if a star schema contains a large fact table spread across twenty nodes,
the joins between the fact table and the dimension tables are most efficient if these
tables are collocated. If all of the tables are in the same database partition group, at
most one dimension table is partitioned correctly for a collocated join. The other
dimension tables cannot be used in a collocated join because the join columns on
the fact table do not correspond to the partitioning key of the fact table.
Consider a table called FACT (C1, C2, C3, ...) partitioned on C1; and a table called
DIM1 (C1, dim1a, dim1b, ...) partitioned on C1; and a table called DIM2 (C2,
dim2a, dim2b, ...) partitioned on C2; and so on.
In this case, you see that the join between FACT and DIM1 is perfect because the
predicate DIM1.C1 = FACT.C1 is collocated. Both of these tables are partitioned on
column C1.
However, the join between DIM2 with the predicate WHERE DIM2.C2 = FACT.C2
cannot be collocated because FACT is partitioned on column C1 and not on
column C2. In this case, you might replicate DIM2 in the database partition group
of the fact table so that the join occurs locally on each partition.
When you create a replicated materialized query table, the source table can be a
single-node table or a multi-node table in a database partition group. In most
cases, the replicated table is small and can be placed in a single-node database
partition group. You can limit the data to be replicated by specifying only a subset
of the columns from the table or by specifying the number of rows through the
predicates used, or by using both methods. The data capture option is not required
for replicated materialized query tables to function.
Indexes on replicated tables are not created automatically. You can create indexes
that are different from those on the source table. However, to prevent constraint
violations that are not present on the source tables, you cannot create unique
indexes or put constraints on the replicated tables. Constraints are disallowed even
if the same constraint occurs on the source table.
Replicated tables can be referenced directly in a query, but you cannot use the
NODENUMBER() predicate with a replicated table to see the table data on a
particular partition.
Use the EXPLAIN facility to see if a replicated materialized query table was used
by the access plan for a query. Whether the access plan chosen by the optimizer
uses the replicated materialized query table depends on the information that needs
to be joined. The optimizer might not use the replicated materialized query table if
the optimizer determines that it would be cheaper to broadcast the original source
table to the other partitions in the database partition group.
Related concepts:
v “Joins” on page 156
Table Queues
Related concepts:
v “Joins” on page 156
v “Join methods” on page 157
v “Join methods in partitioned databases” on page 165
Note: In the diagrams q1, q2, and q3 refer to table queues in the examples. The
tables that are referenced are divided across two database partitions for the
purpose of these scenarios. The arrows indicate the direction in which the
table queues are sent. The coordinator node is partition 0.
A collocated join occurs locally on the partition where the data resides. The
partition sends the data to the other partitions after the join is complete. For the
optimizer to consider a collocated join, the joined tables must be collocated, and all
pairs of the corresponding partitioning key must participate in the equality join
predicates.
q1
q1
Broadcast outer-table joins are a parallel join strategy that can be used if there are
no equality join predicates between the joined tables. It can also be used in other
situations in which it is the most cost-effective join method. For example, a
broadcast outer-table join might occur when there is one very large table and one
very small table, neither of which is partitioned on the join predicate columns.
Instead of partitioning both tables, it might be cheaper to broadcast the smaller
table to the larger table. The following figures provide an example.
q2 q2
• Scan • Scan
LINEITEM LINEITEM
• Apply • Apply
predicates predicates
• Read q2 • Read q2
• Join • Join
• Insert q1 • Insert q1
q1
q1
The ORDERS table is sent to all database partitions that have the LINEITEM table.
Table queue q2 is broadcast to all database partitions of the inner table.
In the directed outer-table join strategy, each row of the outer table is sent to one
partition of the inner table, based on the partitioning attributes of the inner table.
The join occurs on this database partition. The following figure provides an
example.
Scan Scan
LINEITEM LINEITEM
Apply Apply
predicates predicates
Read q2 Read q2
Join Join
Insert into q1 Insert into q1
q1
q1
In the directed inner-table and outer-table join strategy, rows of both the outer and
inner tables are directed to a set of database partitions, based on the values of the
joining columns. The join occurs on these database partitions. The following figure
provides an example. An example is shown in the following figure.
• Read q2 • Read q2
• Read q3 • Read q3
• Join • Join
• Insert q1 • Insert q1
q1
q1
In the broadcast inner-table join strategy, the inner table is broadcast to all the
database partitions of the outer join table. The following figure provides an
example.
• Scan • Scan
q2 LINEITEM LINEITEM
q2
• Apply • Apply
predicates predicates
• Write q3 • Write q3
q3 q3
• Read q2 q3
• Read q2
• Read q3 • Read q3
• Join • Join
• Insert q1 • Insert q1
q1
q1
The LINEITEM table is sent to all database partitions that have the ORDERS table.
Table queue q3 is broadcast to all database partitions of the outer table.
With the directed inner-table join strategy, each row of the inner table is sent to one
database partition of the outer join table, based on the partitioning attributes of the
outer table. The join occurs on this database partition. The following figure
provides an example.
• Scan • Scan
q2 LINEITEM LINEITEM
q2
• Apply • Apply
predicates predicates
• Hash • Hash
ORDERKEY ORDERKEY
q3 q3
• Write q3 • Write q3
• Read q2 q3
• Read q2
• Read q3 • Read q3
• Join • Join
• Insert q1 • Insert q1
q1
q1
Related concepts:
v “Joins” on page 156
v “Join methods” on page 157
v “Join strategies in partitioned databases” on page 164
If the final sorted list of data can be read in a single sequential pass, the results can
be piped. Piping is quicker than non-piped ways of communicating the results of
the sort. The optimizer chooses to pipe the results of a sort whenever possible.
In some cases, the optimizer can choose to push down a sort or aggregation
operation to Data Management Services from the Relational Data Services
component. Pushing down these operations improves performance by allowing the
Data Management Services component to pass data directly to a sort or
aggregation routine. Without this pushdown, Data Management Services first
passes this data to Relational Data Services, which then interfaces with the sort or
aggregation routines. For example, the following query benefits from this
optimization:
SELECT WORKDEPT, AVG(SALARY) AS AVG_DEPT_SALARY
FROM EMPLOYEE
GROUP BY WORKDEPT
When sorting produces the order required for a GROUP BY operation, the
optimizer can perform some or all of the GROUP BY aggregations while doing the
sort. This is advantageous if the number of rows in each group is large. It is even
more advantageous if doing some of the grouping during the sort reduces or
eliminates the need for the sort to spill to disk.
Related concepts:
v “Guidelines for sort performance” on page 236
Related reference:
v “sortheap - Sort heap size” on page 355
v “sheapthres - Sort heap threshold” on page 354
Optimization strategies
This section describes the particular strategies that the optimizer might use for
intra-partition parallelism and multi-dimensional clustering (MDC) tables.
At execution time, multiple database agents called subagents are created to execute
the query. The number of subagents is less than or equal to the degree of
parallelism specified when the SQL statement was compiled.
To parallelize an access plan, the optimizer divides it into a portion that is run by
each subagent and a portion that is run by the coordinating agent. The subagents
pass data through table queues to the coordinating agent or to other subagents. In
a partitioned database, subagents can send or receive data through table queues
from subagents in other database partitions.
Relational scans and index scans can be performed in parallel on the same table or
index. For parallel relational scans, the table is divided into ranges of pages or
rows. A range of pages or rows is assigned to a subagent. A subagent scans its
assigned range and is assigned another range when it has completed its work on
the current range.
For parallel index scans, the index is divided into ranges of records based on index
key values and the number of index entries for a key value. The parallel index
scan proceeds like the parallel table scan with subagents being assigned a range of
records. A subagent is assigned a new range when it has complete its work on the
current range.
The optimizer determines the scan unit (either a page or a row) and the scan
granularity.
Parallel scans provide an even distribution of work among the subagents. The goal
of a parallel scan is to balance the load among the subagents and keep them
equally busy. If the number of busy subagents equals the number of available
processors and the disks are not overworked with I/O requests, then the machine
resources are being used effectively.
The optimizer can choose one of the following parallel sort strategies:
v Round-robin sort
This is also known as a redistribution sort. This method uses shared memory
efficiently redistribute the data as evenly as possible to all subagents. It uses a
round-robin algorithm to provide the even distribution. It first creates an
individual sort for each subagent. During the insert phase, subagents insert into
each of the individual sorts in a round-robin fashion to achieve a more even
distribution of data.
v Partitioned sort
This is similar to the round-robin sort in that a sort is created for each subagent.
The subagents apply a hash function to the sort columns to determine into
which sort a row should be inserted. For example, if the inner and outer tables
of a merge join are a partitioned sort, a subagent can use merge join to join the
corresponding partitions and execute in parallel.
v Replicated sort
This sort is used if each subagent requires all of the sort output. One sort is
created and subagents are synchronized as rows are inserted into the sort. When
the sort is completed, each subagent reads the entire sort. If the number of rows
is small, this sort may be used to rebalance the data stream.
v Shared sort
This sort is the same as a replicated sort, except the subagents open a parallel
scan on the sorted result to distribute the data among the subagents in a way
similar to the round-robin sort.
Subagents can cooperate to produce a temporary table by inserting rows into the
same table. This is called a shared temporary table. The subagents can open
private scans or parallel scans on the shared temporary table depending on
whether the data stream is to be replicated or partitioned.
Otherwise, the subagent can perform a partial aggregation and use another
strategy to complete the aggregation. Some of these strategies are:
v Send the partially aggregated data to the coordinator agent through a merging
table queue. The coordinator completes the aggregation.
v Insert the partially aggregated data into a partitioned sort. The sort is partitioned
on the grouping columns and guarantees that all rows for a set of grouping
columns are contained in one sort partition.
Related concepts:
v “Parallel processing for applications” on page 88
v “Optimization strategies for MDC tables” on page 175
Note: MDC table optimization strategies can also implement the performance
advantages of intra-partition parallelism and inter-partition parallelism.
Consider the following simple example for an MDC table named sales with
dimensions defined on the region and month columns:
SELECT * FROM SALES
WHERE MONTH=’March’ AND REGION=’SE’
For this query, the optimizer can perform a dimension block index lookup to find
blocks in which the month of March and the SE region occur. Then it can quickly
scan only the resulting blocks of the table to fetch the result set.
Related concepts:
v “Table and index management for MDC tables” on page 21
Knowledge of MQTs is integrated into the SQL compiler. In the SQL compiler, the
query rewrite phase and the optimizer match queries with MQTs and determine
whether to substitute an MQT for a query that accesses the base tables. If an MQT
is used, the EXPLAIN facility can provide information about which MQT was
selected.
Because MQTs behave like regular tables in many ways, the same guidelines for
optimizing data access using table space definitions, creating indexes, and issuing
RUNSTATS apply to MQTs.
To help you understand the power of MQTs, the following example shows a
multidimensional analysis query and how it takes advantage of MQTs.
An MQT is created with the sum and count of sales for each level of the following
hierarchies:
v Product
v Location
v Time, composed of year, month, day.
Many queries can be satisfied from this stored aggregate data. The following
example shows how to create an MQT that computes sum and count of sales along
the product group and line dimensions; along the city, state, and country
dimension; and along the time dimension. It also includes several other columns in
its GROUP BY clause.
CREATE TABLE dba.PG_SALESSUM
AS (
SELECT [Link] AS prodline, [Link] AS pgroup,
[Link], [Link], [Link],
Queries that can take advantage of such pre-computed sums would include the
following:
v Sales by month and product group
v Total sales for years after 1990
v Sales for 1995 or 1996
v Sum of sales for a product group or product line
v Sum of sales for a specific product group or product line AND for 1995, 1996
v Sum of sales for a specific country.
While the precise answer is not included in the MQT for any of these queries, the
cost of computing the answer using the MQT could be significantly less than using
a large base table, because a portion of the answer is already computed. MQTs can
reduce expensive joins, sorts, and aggregation of base data.
The first example returns the total sales for 1995 and 1996:
SET CURRENT REFRESH AGE=ANY
The second example returns the total sales by product group for 1995 and 1996:
SET CURRENT REFRESH AGE=ANY
The larger the base tables are, the larger the improvements in response time can be
because the MQT grows more slowly than the base table. MQTs can effectively
eliminate overlapping work among queries by doing the computation once when
the MQTs are built and refreshed and reusing their content for many queries.
Related concepts:
v “The Design Advisor” on page 201
v “Replicated materialized-query tables in partitioned databases” on page 162
Note: Although the DB2® SQL compiler has much information about data source
SQL support, this data may need adjustment over time because data sources
can be upgraded and/or customized. In such cases, make enhancements
known to DB2 by changing local catalog information. Use DB2 DDL
statements (such as CREATE FUNCTION MAPPING and ALTER SERVER)
to update the catalog.
If functions cannot be pushed down to the remote data source, they can
significantly impact query performance. Consider the effect of forcing a selective
predicate to be evaluated locally instead of at the data source. Such evaluation
could require DB2 to retrieve the entire table from the remote data source and then
filter it locally against the predicate. Network constraints and large table size could
cause performance to suffer.
Operators that are not pushed down can also significantly affect query
performance. For example, having a GROUP BY operator aggregate remote data
locally could also require DB2 to retrieve the entire table from the remote data
source.
For example, assume that a nickname N1 references the data source table
EMPLOYEE in a DB2 for OS/390® or z/OS data source. Also assume that the table
has 10,000 rows, that one of the columns contains the last names of employees, and
that one of the columns contains salaries. Consider the following statement:
SELECT LASTNAME, COUNT(*) FROM N1
WHERE LASTNAME > ’B’ AND SALARY > 50000
GROUP BY LASTNAME;
In general, the goal is to ensure that the optimizer evaluates functions and
operators on data sources. Many factors affect whether a function or an SQL
operator is evaluated at a remote data source. Factors to be evaluated are classified
in the following three groups:
v Server characteristics
v Nickname characteristics
v Query characteristics
SQL Capabilities: Each data source supports a variation of the SQL dialect and
different levels of functionality. For example, consider the GROUP BY list. Most
data sources support the GROUP BY operator, but some limit the number of items
on the GROUP BY list or have restrictions on whether an expression is allowed on
the GROUP BY list. If there is a restriction at the remote data source, DB2 might
have to perform the GROUP BY operation locally.
SQL Restrictions: Each data source might have different SQL restrictions. For
example, some data sources require parameter markers to bind values to remote
SQL statements. Therefore, parameter marker restrictions must be checked to
ensure that each data source can support such a bind mechanism. If DB2 cannot
determine a good method to bind a value for a function, this function must be
evaluated locally.
SQL Limits: Although DB2 might allow the use of larger integers than its remote
data sources, values that exceed remote limits cannot be embedded in statements
sent to data sources. Therefore, the function or operator that operates on this
constant must be evaluated locally.
Server Specifics: Several factors fall into this category. One example is whether
NULL values are sorted as the highest or lowest value, or depend on the ordering.
If NULL values are sorted at a data source differently from DB2, ORDER BY
operations on a nullable expression cannot be remotely evaluated.
The following operations might be pushed down if collating sequences are the
same:
v Comparisons of character or numeric data
v Character range comparison predicates
v Sorts
You might get unusual results, however, if the weighting of null characters is
different between the federated database and the data source. Comparison
statements might return unexpected results if you submit statements to a
case-insensitive data source. The weights assigned to the characters ″I″ and ″i″ in a
case-insensitive data source are the same. DB2, by default, is case sensitive and
assigns different weights to the characters.
To improve performance, the federated server allows sorts and comparisons to take
place at data sources. For example, in DB2 UDB for OS/390 or z/OS, sorts defined
by ORDER BY clauses are implemented by a collating sequence based on an
EBCDIC code page. To use the federated server to retrieve DB2 for OS/390 or
z/OS data sorted in accordance with ORDER BY clauses, configure the federated
database so that it uses a predefined collating sequence based on the EBCDIC code
page.
If the collating sequences of the federated database and the data source differ, DB2
retrieves the data to the federated database. Because users expect to see the query
results ordered by the collating sequence defined for the federated server, by
ordering the data locally the federated server ensures that this expectation is
fulfilled. Submit your query in pass-through mode, or define the query in a data
source view if you need to see the data ordered in the collating sequence of the
data source.
DB2 Type Mapping and Function Mapping Factors: The default local data type
mappings provided by DB2 are designed to provide sufficient buffer space for each
data source data type, which avoids loss of data. Users can customize the type
mapping for a specific data source to suit specific applications. For example, if you
are accessing an Oracle data source column with a DATE data type, which by
default is mapped to the DB2 TIMESTAMP data type, you might change the local
data type to the DB2 DATE data type.
In the following three cases, DB2 can compensate for functions that a data source
does not support:
v The function does not exist at the remote data source.
v The function exists, but the characteristics of the operand violate function
restrictions. An example of this situation is the IS NULL relational operator.
Most data sources support it, but some may have restrictions, such as only
allowing a column name on the left hand side of the IS NULL operator.
Local data type of a nickname column: Ensure that the local data type of a
column does not prevent a predicate from being evaluated at the data source. Use
the default data type mappings to avoid possible overflow. However, a joining
predicate between two columns of different lengths might not be considered at the
data source whose joining column is shorter, depending on how DB2 binds the
longer column. This situation can affect the number of possibilities that the DB2
optimizer can evaluate in a joining sequence. For example, Oracle data source
columns created using the INTEGER or INT data type are given the type
NUMBER(38). A nickname column for this Oracle data type is given the local data
type FLOAT because the range of a DB2 integer is from 2**31 to (-2**31)-1, which is
roughly equal to NUMBER(9). In this case, joins between a DB2 integer column
and an Oracle integer column cannot take place at the DB2 data source (shorter
joining column); however, if the domain of this Oracle integer column can be
accommodated by the DB2 INTEGER data type, change its local data type with the
ALTER NICKNAME statement so that the join can take place at the DB2 data
source.
Column Options: Use the SQL statement ALTER NICKNAME to add or change
column options for nicknames.
Use the numeric_string option to indicate whether the values in that column are
always numbers without trailing blanks.
If you set numeric_string to ‘Y’ for a column, you are informing the
optimizer that this column contains no blanks that could interfere with
sorting of the column data. This option is useful when the collating
sequence of a data source is different from DB2. Columns marked with
this option are not excluded from local (data source) evaluation because
of a different collating sequence.
varchar_no_trailing_blanks Specifies whether this data source uses non-blank padded VARCHAR ‘N‘
comparison semantics. For variable-length character strings that contain
no trailing blanks, non-blank-padded comparison semantics of some
DBMSs return the same results as DB2 comparison semantics. If you are
certain that all VARCHAR table/view columns at a data source contain
no trailing blanks, consider setting this server option to ’Y’ for a data
source. This option is often used with Oracle data sources. Ensure that
you consider all objects that might have nicknames, including views.
’Y’ This data source has non-blank-padded comparison semantics
similar to DB2.
’N’ This data source does not have the same non-blank-padded
comparison semantics as DB2.
A query can reference an SQL operator that might involve nicknames from
multiple data sources. The operation must take place at DB2 to combine the results
from two referenced data sources that use one operator, such as a set operator (e.g.
UNION). The operator cannot be evaluated at a remote data source directly.
Related concepts:
v “Guidelines for analyzing where a federated query is evaluated” on page 182
Consider the following key questions when you investigate ways to increase
pushdown opportunities:
v Why isn’t this predicate being evaluated remotely?
This question arises when a predicate is very selective and thus could be used to
filter rows and reduce network traffic. Remote predicate evaluation also affects
whether a join between two tables of the same data source can be evaluated
remotely.
Areas to examine include:
– Subquery predicates. Does this predicate contain a subquery that pertains to
another data source? Does this predicate contain a subquery involving an
SQL operator that is not supported by this data source? Not all data sources
support set operators in a subquery predicate.
– Predicate functions. Does this predicate contain a function that cannot be
evaluated by this remote data source? Relational operators are classified as
functions.
– Predicate bind requirements. Does this predicate, if remotely evaluated,
require bind-in of some value? If so, would it violate SQL restrictions at this
data source?
– Global optimization. The optimizer may have decided that local processing is
more cost effective.
v Why isn’t the GROUP BY operator evaluated remotely?
There are several areas you can check:
– Is the input to the GROUP BY operator evaluated remotely? If the answer is
no, examine the input.
– Does the data source have any restrictions on this operator? Examples
include:
- Limited number of GROUP BY items
- Limited byte counts of combined GROUP BY items
- Column specification only on the GROUP BY list
– Does the data source support this SQL operator?
– Global optimization. The optimizer may have decided that local processing is
more cost effective.
– Does the GROUP BY operator clause contain a character expression? If it
does, verify that the remote data source has the same case sensitivity as DB2.
v Why isn’t the set operator evaluated remotely?
There are several areas you can check:
– Are both of its operands completely evaluated at the same remote data
source? If the answer is no and it should be yes, examine each operand.
– Does the data source have any restrictions on this set operator? For example,
are large objects or long fields valid input for this specific set operator?
v Why isn’t the ORDER BY operation evaluated remotely?
Consider:
Related concepts:
v “Character-conversion guidelines” on page 86
v “Global analysis of federated database queries” on page 186
v “Remote SQL generation and global optimization in federated databases” on
page 184
v “Federated query information” on page 573
The optimizer uses the output of pushdown analysis to decide whether each
operation is evaluated locally at DB2® or remotely at a data source. It bases its
decision on the output of its cost model, which includes not only the cost of
evaluating the operation but also the cost of transmitting the data or messages
between DB2 and data sources.
Although the goal is to produce an optimized query, the following major factors
affect the output from global optimization and thus affect query performance.
v Server characteristics
v Nickname characteristics
The following data source server factors can affect global optimization:
v Relative ratio of CPU speed
Use the cpu_ratio server option to specify how fast or slow the data-source CPU
speed is compared with the DB2 CPU. A low ratio indicates that the data-source
computer CPU is faster than the DB2 computer CPU. If the ratio is low, the DB2
optimizer is more likely to consider pushing down CPU-intensive operations to
the data source.
v Relative ratio of I/O speed
Use the io_ratio server option to indicate how much faster or slower the data
source system I/O speed is compared with the DB2 system. A low ratio
indicates that the data source workstation I/O speed is faster than the DB2
workstation I/O speed. If the ratio is low, the DB2 optimizer considers pushing
down I/O-intensive operations to the data source.
v Communication rate between DB2 and the data source
Index considerations: To optimize queries, DB2 can use information about indexes
at data sources. For this reason, it is important that the index information available
to DB2 is current. The index information for nicknames is initially acquired when
the nickname is created. Index information is not gathered for view nicknames.
Before you issue CREATE INDEX statements against a nickname for a view,
consider whether you need one. If the view is a simple SELECT on a table with an
index, creating local indexes on the nickname to match the indexes on the table at
the data source can significantly improve query performance. However, if indexes
are created locally over views that are not simple select statements, such as a view
created by joining two tables, query performance might suffer. For example, you
create an index over a view that is a join of two tables, the optimizer might choose
that view as the inner element in a nested-loop join. The query will have poor
performance because the join is evaluated several times. An alternative is to create
nicknames for each of the tables referenced in the data source view and create a
local view at DB2 that references both nicknames.
Catalog statistics considerations: System catalog statistics describe the overall size
of nicknames and the range of values in associated columns. The optimizer uses
these statistics when it calculates the least-cost path for processing queries that
contain nicknames. Nickname statistics are stored in the same catalog views as
table statistics.
Although DB2 can retrieve the statistical data stored at a data source, it cannot
automatically detect updates to existing statistical data at data sources.
Furthermore, DB2 cannot handle changes in object definition or structural changes,
such as adding a column, to objects at data sources. If the statistical data or
structural data for an object has changed, you have two choices:
v Run the equivalent of RUNSTATS at the data source. Then drop the current
nickname and re-create it. Use this approach if structural information has
changed.
v Manually update the statistics in the [Link] view. This approach
requires fewer steps but it does not work if structural information has changed.
Related concepts:
v “Server options affecting federated databases” on page 94
v “Guidelines for analyzing where a federated query is evaluated” on page 182
v “Global analysis of federated database queries” on page 186
Consider the following optimization questions and key areas to investigate for
performance improvements:
v Why isn’t a join between two nicknames of the same data source being
evaluated remotely?
Areas to examine include:
– Join operations. Can the data source support them?
– Join predicates. Can the join predicate be evaluated at the remote data source?
If the answer is no, examine the join predicate.
– Number of rows in the join result (with Visual Explain). Does the join
produce a much larger set of rows than the two nicknames combined? Do the
numbers make sense? If the answer is no, consider updating the nickname
statistics manually ([Link]).
v Why isn’t the GROUP BY operator being evaluated remotely?
Areas to examine include:
– Operator syntax. Verify that the operator can be evaluated at the remote data
source.
– Number of rows. Check the estimated number of rows in the GROUP BY
operator input and output using visual explain. Are these two numbers very
close? If the answer is yes, the DB2 optimizer considers it more efficient to
evaluate this GROUP BY locally. Also, do these two numbers make sense? If
the answer is no, consider updating the nickname statistics manually
([Link]).
v Why is the statement not being completely evaluated by the remote data source?
The DB2 optimizer performs cost-based optimization. Even if pushdown analysis
indicates that every operator can be evaluated at the remote data source, the
optimizer still relies on its cost estimate to generate a globally optimal plan.
There are a great many factors that can contribute to that plan. For example,
even though the remote data source can process every operation in the original
query, its CPU speed is much slower than the CPU speed for DB2 and thus it
may turn out to be more beneficial to perform the operations at DB2 instead. If
results are not satisfactory, verify your server statistics in
[Link].
v Why does a plan generated by the optimizer, and completely evaluated at a
remote data source, have much worse performance than the original query
executed directly at the remote data source?
Areas to examine include:
– The remote SQL statement generated by the DB2 optimizer. Ensure that it is
identical to the original query. Check for predicate ordering changes. A good
query optimizer should not be sensitive to the predicate ordering of a query;
unfortunately, not all DBMS optimizers are identical, and thus it is likely that
the optimizer of the remote data source may generate a different plan based
on the input predicate ordering. If this is true, this is a problem inherent in
the remote optimizer. Consider either modifying the predicate ordering on the
input to DB2 or contacting the service organization of the remote data source
for assistance.
Also, check for predicate replacements. A good query optimizer should not be
sensitive to equivalent predicate replacements; unfortunately, not all DBMS
optimizers are identical, and thus it is possible that the optimizer of the
Related concepts:
v “Server options affecting federated databases” on page 94
v “Federated database pushdown analysis” on page 178
v “Guidelines for analyzing where a federated query is evaluated” on page 182
v “Remote SQL generation and global optimization in federated databases” on
page 184
v “Federated query information” on page 573
v “Example five: federated database plan” on page 584
You collect and use explain data for the following reasons:
v To understand how the database manager accesses tables and indexes to satisfy
your query
v To evaluate your performance-tuning actions
When you change some aspect of the database manager, the SQL statements, or
the database, you should examine the explain data to find out how your action
has changed performance.
Before you can capture explain information, you create the relational tables in
which the optimizer stores the explain information and you set the special registers
that determine what kind of explain information is captured.
To display explain information, you can use either a command-line tool or Visual
Explain. The tool that you use determines how you set the registry variables that
determine what explain data is collected. For example, if you expect to use Visual
Explain only, you need only capture snapshot information. If you expect to
perform detailed analysis with one of the command-line utilities or with custom
SQL statements against the explain tables, you should capture all explain
information.
Related concepts:
v “Explain tools” on page 190
v “Guidelines for using explain information” on page 191
v “The explain tables and organization of explain information” on page 193
v “Guidelines for capturing explain information” on page 198
v “Guidelines for analyzing explain information” on page 200
v “SQL explain tools” on page 551
Explain tools
DB2® provides a comprehensive explain facility that provides detailed information
about the access plan that the optimizer chooses for an SQL statement. The tables
that store explain data are accessible on all supported platforms and contain
information for both static and dynamic SQL statements. Several tools or methods
give you the flexibility you need to capture, display, and analyze explain
information.
Detailed optimizer information that allows for in-depth analysis of an access plan
is stored in explain tables separate from the actual access plan itself. Use one or
more of the following methods of getting information from the explain tables:
v Use Visual Explain to view explain snapshot information
Invoke Visual Explain from the Control Center to see a graphical display of a
query access plan. You can analyze both static and dynamic SQL statements.
Visual Explain allows you to view snapshots captured or taken on another
platform. For example, a Windows® NT client can graph snapshots generated on
a DB2 for HP-UX server. To do this, both of the platforms must be at a Version 5
level or later.
v Use the db2exfmt tool to display explain information in preformatted output.
v Use the db2expln and dynexpln tools
To see the access plan information available for one or more packages of static
SQL statements, use the db2expln tool from the command line. db2expln shows
the actual implementation of the chosen access plan. It does not show optimizer
information.
The dynexpln tool, which uses db2expln within it, provides a quick way to
explain dynamic SQL statements that contain no parameter markers. This use of
db2expln from within dynexpln is done by transforming the input SQL statement
into a static statement within a pseudo-package. When this occurs, the
information may not always be completely accurate. If complete accuracy is
desired, use the explain facility.
The db2expln tool does provide a relatively compact and English-like overview of
what operations will occur at run-time by examining the actual access plan
generated.
v Write your own queries against the explain tables
Writing your own queries allows for easy manipulation of the output and for
comparison among different queries or for comparisons of the same query over
time.
Note: The location of the command-line explain tools and others, such as db2batch,
dynexpln, db2vexp , and db2_all, is in the misc subdirectory of the sqllib
directory. If the tools are moved from this path, the command-line methods
might not work.
Related concepts:
v “dynexpln” on page 558
v “Description of db2expln and dynexpln output” on page 558
v “Examples of db2expln and dynexpln output” on page 576
Related reference:
v Appendix D, “db2exfmt - Explain Table Format,” on page 587
v “db2expln - SQL Explain” on page 552
To help you understand the reasons for changes in query performance, you need
the before and after explain information which you can obtain by performing the
following steps:
v Capture explain information for the query before you make any changes and
save the resulting explain tables, or you might save the output from the
db2exfmt explain tool.
v Save or print the current catalog statistics if you do not want to, or cannot,
access Visual Explain to view this information. You might also use the db2look
productivity tool to help perform this task.
The information that you collect in this way provides a reference point for future
analysis. For dynamic SQL statements, you can collect this information when you
run your application for the first time. For static SQL statements, you can also
collect this information at bind time. To analyze a performance change, you
compare the information that you collected with information that you collect about
the query and environment when you start your analysis.
As a simple example, your analysis might show that an index is no longer being
used as part of the access path. Using the catalog statistics information in Visual
Explain, you might notice that the number of index levels (NLEVELS column) is
now substantially higher than when the query was first bound to the database.
You might then choose to perform one of these actions:
v Reorganize the index
v Collect new statistics for your table and indexes
v Gather explain information when rebinding your query.
After you perform one of the actions, examine the access plan again. If the index is
used again, performance of the query might no longer be a problem. If the index is
still not used or if performance is still a problem, perform a second action and
examine the results. Repeat these steps until the problem is resolved.
You can take a number of actions to help improve query performance, such as
adjusting configuration parameters, adding containers, collecting fresh catalog
statistics, and so on.
After you make a change in any of these areas, you can use the SQL explain
facility to determine the impact, if any, that the change has on the access plan
chosen. For example, if you add an index or materialized query table (MQT) based
on the index guidelines, the explain data can help you determine whether the
index or materialized query table is actually used as you expected.
Although the explain output provides information that allows you to determine
the access plan that was chosen and its relative cost, the only way to accurately
measure the performance improvement for a query is to use benchmark testing
techniques.
Related concepts:
v “SQL explain facility” on page 189
v “The explain tables and organization of explain information” on page 193
v “Guidelines for capturing explain information” on page 198
v “The Design Advisor” on page 201
v “Materialized query tables” on page 176
Related reference:
v Appendix D, “db2exfmt - Explain Table Format,” on page 587
Explain table information reflects the relationships between operators and data
objects in the access plan. The following diagram shows the relationships between
these tables.
Note: Not all of the tables above are created by default. To create them, run the
[Link] script found in the misc subdirectory of the sqllib
subdirectory.
Explain tables might be common to more than one user. However, the explain
tables can be defined for one user, and then aliases can be defined for each
additional user using the same name to point to the defined tables. Each user
sharing the common explain tables must have insert permission on those tables.
Related concepts:
v “SQL explain facility” on page 189
v “Explain information for data objects” on page 194
v “Explain information for instances” on page 196
v “Explain information for data operators” on page 195
v “SQL explain tools” on page 551
Object Statistics: The explain facility records information about the object, such as
the following:
v The creation time
v The last time that statistics were collected for the object
v An indication of whether or not the data in the object is ordered (only table or
index objects)
v The number of columns in the object (only table or index objects)
Related concepts:
v “The explain tables and organization of explain information” on page 193
v “Explain information for instances” on page 196
v “Explain information for data operators” on page 195
v “Guidelines for analyzing explain information” on page 200
In addition to showing the operators used in an access plan and information about
each operator, explain information also shows the cumulative effects of the access
plan.
Timerons are an invented relative unit of measure. Timerons are determined by the
optimizer based on internal values such as statistics that change as the database is
used. As a result, the timerons measure for a SQL statement are not guaranteed to
be the same every time the estimated cost in timerons is determined.
Related concepts:
v “The explain tables and organization of explain information” on page 193
v “Explain information for data objects” on page 194
v “Explain information for instances” on page 196
v “Guidelines for analyzing explain information” on page 200
Cost Estimation: For each explained statement, the optimizer records an estimate
of the relative cost of executing the chosen access plan. This cost is stated in an
invented relative unit of measure called a timeron. No estimate of elapsed times is
provided, for the following reasons:
v The SQL optimizer does not estimate elapsed time but only resource
consumption.
v The optimizer does not model all factors that can affect elapsed time. It ignores
factors that do not affect the efficiency of the access plan. A number of runtime
factors affect the elapsed time, including the system workload, the amount of
resource contention, the amount of parallel processing and I/O, the cost of
returning rows to the user, and the communication time between the client and
server.
Statement Text: Two versions of the text of the SQL statement are recorded for
each statement explained. One version is the code that the SQL compiler receives
from the application. The other version is reverse-translated from the internal
compiler representation of the query. Although this translation looks similar to
other SQL statements, it does not necessarily follow correct SQL syntax nor does it
necessarily reflect the actual content of the internal representation as a whole. This
translation is provided only to allow you to understand the SQL context in which
the SQL optimizer chose the access plan. To understand how the SQL compiler has
rewritten your query for better optimization, compare the user-written statement
text to the internal representation of the SQL statement. The rewritten statement
also shows you other elements in the environment affecting your statement, such
as triggers and constraints. Some keywords used by this “optimized” text are:
$Cn The name of a derived column, where n represents
an integer value.
$CONSTRAINT$ The tag used to indicate the name of a constraint
added to the original SQL statement during
compilation. Seen in conjunction with the
$WITH_CONTEXT$ prefix.
$[Link] The name of a derived table, where n represents an
integer value.
$INTERNAL_FUNC$ The tag indicates the presence of a function used
Related concepts:
v “The explain tables and organization of explain information” on page 193
v “Explain information for data objects” on page 194
v “Explain information for data operators” on page 195
v “Guidelines for analyzing explain information” on page 200
Related reference:
v “comm_bandwidth - Communications bandwidth” on page 456
v “sortheap - Sort heap size” on page 355
v “locklist - Maximum storage for lock list” on page 340
v “maxlocks - Maximum percent of lock list before escalation” on page 369
v “dbheap - Database heap” on page 339
v “cpuspeed - CPU speed” on page 457
v “avg_appls - Average number of active applications” on page 378
v “dft_degree - Default degree” on page 431
Related concepts:
v “SQL explain facility” on page 189
v “Guidelines for using explain information” on page 191
v “The explain tables and organization of explain information” on page 193
v “The Design Advisor” on page 201
v “Guidelines for analyzing explain information” on page 200
v “SQL explain tools” on page 551
Related concepts:
v “SQL explain facility” on page 189
v “SQL explain tools” on page 551
v “Description of db2expln and dynexpln output” on page 558
| You can have the Design Advisor implement some or all of these recommendations
| immediately or schedule them for a later time.
| Using either the Design Advisor GUI or the command-line tool, the Design
| Advisor can help simplify the following tasks:
| Planning for or setting up a new database
| While designing your database use the Design Advisor to:
| v Generate design alternatives in a test environment of a partitioned
| database environment, and of indexes, MQTs, and MDC tables.
| v For partitioned database environments, you can use the Design Advisor
| to:
| – Determine the partitioning strategy before loading data into a
| database.
| – Assist in migrating from a single-partition DB2 database to a
| multiple-partition DB2 database.
| – Assist in migrating from another database product to a
| multiple-partition DB2 database.
| v Evaluate indexes, MQTs, MDC tables, or partitioning strategies that have
| been generated manually.
| Workload performance tuning
| After your database is set up, you can use the Design Advisor to:
| v Improve performance of a particular statement or workload.
| v Improve general database performance, using the performance of a
| sample workload as a gauge.
| v Improve performance of the most frequently executed queries, for
| example, as identified by the Activity Monitor.
| v Determine how to optimize the performance of a new key query.
| v Respond to Health Center recommendations regarding shared memory
| utility or sort heap problems in a sort-intensive workload.
| v Find objects that are not used in a workload.
| The ADVISE_INSTANCE table is also updated with one row each time that the
| Design Advisor runs:
| v The START_TIME field and the END_TIME field show the start and stop times
| of the utility, respectively.
| v The STATUS field will contain 'COMPLETED' if the utility ended successfully.
| v The MODE field indicates whether the -m option was used.
| v The COMPRESSION field indicates the type of compression used.
| You can save the Design Advisor recommendations to a file using the -o option.
| The saved Design Advisor output consists of the following elements:
| v CREATE STATEMENTS given for new indexes, MQTs, partitioning strategies,
| and MDC tables.
| v REFRESH statements for MQTs.
| v RUNSTATS commands for new objects.
| v Existing MQTs and indexes will appear in the recommended script if they were
| and are used to execute the workload.
| Note: The COLSTATS column of the ADVISE_MQT table contains the column
| statistics for an MQT. The statistics are in an XML structure as follows:
| <?xml version=\"1.0\" encoding=\"USASCII\"?>
| <colstats>
| <column>
| <name>COLNAME1</name>
| <colcard>1000</colcard>
| <high2key>999</high2key>
| <low2key>2</low2key>
| </column>
| ....
|
| <column>
| <name>COLNAME100</name>
| <colcard>55000</colcard>
| <high2key>49999</high2key>
| <low2key>100</low2key>
| </column>
| </colstats>
| Note that the XML structure can contain more than one column. For each
| column, the column cardinality (that is, the number of values in the
| column) is shown, and optionally, the high2 and low2 keys.
| After some minor modifications, you can run this output file as a CLP script to
| create the recommended objects. The modifications that you might want to
| perform include:
| v Combining all of the RUNSTATS command statements into a single RUNSTATS
| invocation on the new or modified objects.
| v Providing more usable object names than the system-generated IDs.
| v Removing or commenting out any DDL for objects that you do not want to
| implement immediately.
| Related reference:
| v “db2advis - DB2 Design Advisor Command” in the Command Reference
| Procedure:
| From the Design Advisor GUI workload page, you can create a new workload file,
| or modify a previously existing workload file. You can import statements into the
| file from several sources:
| v A delimited text file
| v An Event Monitor table
| v Query Patroller historical data tables by using the -qp option from the command
| line
| v Explained statements in the EXPLAINED_STATEMENT table
| v Recent SQL statements that have been captured with a DB2 snapshot.
| After you import your SQL statements, you can add, change, modify, or remove
| statements and modify their frequency.
| To run the Design Advisor on a set of SQL statements contained in a workload file:
| 1. Create a workload file manually, separating each SQL statement with a
| semicolon, or import SQL statements from one or more of the sources listed
| above.
| 2. Set the frequency of the statements in the workload. Every statement in the
| workload file is assigned a frequency of 1 by default. The frequency of an SQL
| statement represents the number of times the statement occurs within a
| workload relative to the number of times that other statements occur. For
| example, a particular SELECT statement might occur 100 times in a workload,
| while another SELECT statement occurs 10 times. To represent the relative
| frequency of these two statements, you could assign the first SELECT statement
| a frequency of 10, while the second select statement has a frequency of 1. You
| can manually change the frequency or weight that a particular statement has in
| the workload by inserting the following line after the statement - - # SET
| FREQUENCY n where n is the frequency value that you want to assign to the
| statement.
| 3. Run the db2advis command using the -i option followed by the name of the
| workload file.
| Related concepts:
| v “The Design Advisor” on page 201
| Procedure:
| 1. Update the product license key for DB2 UDB ESE.
| 2. Create at least one table space in a multiple-partition database partition group.
| Related tasks:
| v “Registering the DB2 product license key using the db2licm command” in the
| Installation and Configuration Supplement
| Related concepts:
| v “The Design Advisor” on page 201
Memory usage
This section describes how the database manager uses memory and lists the
parameters that control the database manager and database use of memory.
The figure below shows different portions of memory that the database manager
allocates for various uses.
Note: This figure does not show how memory is used in an Enterprise Server
Edition environment, which comprises multiple logical nodes. In such an
environment, each node contains a Database Manager Shared Memory set.
Memory is allocated for each instance of the database manager when the following
events occur:
v When the database manager is started (db2start): Database manager global
shared memory is allocated and remains allocated until the database manager is
stopped (db2stop). This area contains information that the database manager
uses to manage activity across all database connections. When the first
application connects to a database, both global and private memory areas are
allocated.
v When a database is activated or connected to for the first time: Database global
memory is allocated. Database global memory is used across all applications that
might connect to the database. The size of the database global memory is
specified by the database_memory configuration parameter. You can specify more
memory than is needed initially so that the additional memory can be
dynamically distributed later. Although the total amount of database global
memory cannot be increased or decreased while the database is active, memory
for areas contained in database global memory can be adjusted. Such areas
include the buffer pools, the lock list, the database heap and utility heap, and
the package cache, and the catalog cache. In an environment in which the
database manager intra-partition parallelism configuration parameter
(intra_parallel) is enabled, or in an environment in which the connection
concentrator is enabled, the shared sort heap is also allocated as part of the
database global memory.
v When an application connects to a database: In a partitioned database
environment, in a non-partitioned database with the database manager
intra-partition parallelism configuration parameter (intra_parallel) enabled, or in
an environment in which the connection concentrator is enabled, multiple
applications can be assigned to application groups to share memory. Each
application group has its own allocation of shared memory. In the
application-group shared memory, each application has its own application
control heap but uses the share heap of the application group.
The following three database configuration parameters determine the size of the
application group memory:
– The appgroup_mem_sz parameter, which specifies the size of the shared
memory for the application group
– The groupheap_ratio parameter, which specifies the percent of the
application-group shared memory allowed for the shared heap
212 Administration Guide: Performance
– The app_ctl_heap_sz parameter, which specifies the size of the control heap for
each application in the group.
The performance advantage of grouping application memory use is improved
cache and memory-use efficiency.
Some elements of application global memory can also be resized dynamically.
v When an agent is created: This event is not shown in the figure. Agent private
memory is allocated for an agent when the agent is assigned as the result of a
connect request or a new SQL request in a parallel environment, Agent private
memory is allocated for the agent and contains memory allocations that is used
only by this specific agent, such as the sort heap and the application heap.
When a database is already in use by one application, only agent private
memory and application global shared memory is allocated for subsequent
connecting applications.
The figure also lists the following configuration parameter settings, which limit the
amount of memory that is allocated for each specific purposes. Note that in a
partitioned database environment, this memory is allocated on each database
partition.
v numdb
This parameter specifies the maximum number of concurrent active databases
that different applications can use. Because each database has its own global
memory area, the amount of memory that might be allocated increases if you
increase the value of this parameter.
v maxappls
This parameter specifies the maximum number of applications that can
simultaneously connect to a single database. It affects the amount of memory
that might be allocated for agent private memory and application global
memory for that database. Note that this parameter can be set differently for
every database.
v maxagents and max_coordagents for parallel processing
These parameters are not shown in the figure. They limit the number of
database manager agents that can exist simultaneously across all active
databases in an instance. Together with maxappls, these parameters limit the
amount of memory allocated for agent private memory and application global
memory.
Related concepts:
v “Database manager shared memory” on page 213
v “The FCM buffer pool and memory requirements” on page 215
v “Global memory and parameters that control it” on page 216
v “Guidelines for tuning parameters that affect memory usage” on page 218
v “Memory management” on page 32
The following figure shows how memory is used to support applications. The
configuration parameters shown allow you to control the size of this memory, by
Package cache
(pckcachesz)
(app_ctl_heap_sz)
Agent/Application Note: Box size does not indicate relative size of memory.
shared memory
Application support
layer heap (aslheapsz)
You can predict and control the size of this space by reviewing information about
database agents. Agents running on behalf of applications require substantial
For partitioned database systems, the fast communications manager (FCM) requires
substantial memory space, especially if the value of fcm_num_buffers is large. In
addition, the FCM memory requirements are either allocated from the FCM Buffer
Pool, or from both the Database Manager Shared Memory and the FCM Buffer
Pool, depending on whether or not the partitioned database system uses multiple
logical nodes.
Related concepts:
v “Organization of memory use” on page 211
v “Global memory and parameters that control it” on page 216
v “Buffer pool management” on page 220
v “Guidelines for tuning parameters that affect memory usage” on page 218
v “Memory management” on page 32
Legend
Figure 21. FCM buffer pool when multiple logical nodes are not used
If you have a partitioned database system that uses multiple logical nodes, the
Database Manager Shared Memory and FCM Buffer Pool are as shown below.
Legend
Figure 22. FCM buffer pool when multiple logical nodes are used
For configuring the fast communications manager (FCM), start with the default
value for the number of FCM Buffers (fcm_num_buffers). For more information
about FCM on AIX® platforms, refer to the description of the DB2_FORCE_FCP_BP
registry variable.
To tune this parameter, use the database system monitor to monitor the low water
mark for the free buffers.
Related concepts:
v “Database manager shared memory” on page 213
Related reference:
v “estore_seg_sz - Extended storage memory segment size” on page 373
v “max_coordagents - Maximum number of coordinating agents” on page 379
v “num_estore_segs - Number of extended storage memory segments” on page
373
v “sortheap - Sort heap size” on page 355
v “maxagents - Maximum number of agents” on page 380
Some UNIX® operating systems allocate swap space when a process allocates
memory and not when a process is paged out to swap space. For these systems,
make sure that you provide as much paging space as total shared memory space.
Note: To change the size of a buffer pool, use the DDL statement, ALTER
BUFFERPOOL.
Notes:
v Benchmark tests provide the best information about setting appropriate
values for memory parameters. In benchmarking, typical and worst-case
SQL statements are run against the server and the values of the
parameters are modified until the point of diminishing return for
performance is found. If performance versus parameter values is
graphed, the point at which the curve begins to plateau or decline
indicates the point at which additional allocation provides no additional
value to the application and is therefore simply wasting memory.
v The upper limits of memory allocation for several parameters may be
beyond the memory capabilities of existing hardware and operating
systems. These limits allow for future growth.
v For valid parameter ranges, refer to the detailed information about each
parameter.
Related concepts:
v “Organization of memory use” on page 211
v “Database manager shared memory” on page 213
v “Global memory and parameters that control it” on page 216
Related reference:
v “app_ctl_heap_sz - Application control heap size” on page 346
v “fcm_num_buffers - Number of FCM buffers” on page 444
v “sheapthres - Sort heap threshold” on page 354
v “aslheapsz - Application support layer heap size” on page 358
Buffer pools
Buffer pools are a critically important memory component. This section describes
buffer pools and provides information about managing them for good
performance.
When an application accesses a row of a table for the first time, the database
manager places the page containing that row in the buffer pool. The next time any
application requests data, the database manager looks for it in the buffer pool. If
the requested data is in the buffer pool, it can be retrieved without disk access,
resulting in faster performance.
Memory is allocated for the buffer pool when a database is activated or when the
first application connects to the database. Buffer pools can also be created,
dropped, and resized while the database is manager is running. If you use the
IMMEDIATE keyword when you use ALTER BUFFERPOOL to increase the size of
the buffer pool, memory is allocated as soon as you enter the command if the
memory is available. If the memory is not available, the changed occurs when all
applications are disconnected and the database is reactivated. If you decrease the
size of the buffer pool, memory is deallocated at commit time. When all
applications are disconnected, the buffer pool memory is de-allocated.
Note: To reduce the necessity of increasing the size of the dbheap database
configuration parameter when buffer-pool sizes increase, nearly all
buffer-pool memory, which includes page descriptors, buffer-pool
descriptors, and the hash tables, comes out of the database shared memory
set and is sized automatically.
Pages remain in the buffer pool until the database is shut down, or until the space
occupied by a page is required for another page. The following criteria determine
which page is removed to bring in another page:
v How recently the page was referenced
v The probability that the page will be referenced again by the last agent that
looked at it
v The type of data on the page
v Whether the page was changed in memory but not written out to disk (Changed
pages are always written to disk before being overwritten.)
In order for pages to be accessed from memory again, changed pages are not
removed from the buffer pool after they are written out to disk unless the space is
needed.
When you create a buffer pool, the default page size is 4 KB but you can specify a
page size of 4 KB, 8 KB, 16 KB, or 32 KB. Because pages can be read into a buffer
pool only if the table-space page size is the same as the buffer-pool page size, the
page size of your table spaces should determine the page size that you specify for
buffer pools. You cannot alter the page size of the buffer pool after you create it.
You must create a new buffer pool with a different page size.
Note: On 32-bit platforms that run Windows® NT, you can create large buffer
pools if you have enabled Address Windowing Extensions (AWE) or
Advanced Server and Data Center Server on Windows 2000.
Related concepts:
v “Organization of memory use” on page 211
v “Secondary buffer pools in extended memory on 32-bit platforms” on page 221
v “Buffer pool management of data pages” on page 223
v “Illustration of buffer pool data-page management” on page 225
v “Management of multiple database buffer pools” on page 226
If you define some of the real addressable memory as an extended storage cache,
this memory can no longer be used for other purposes, such as a JFS-cache or as
process private address space. More system paging might occur if you allocate real
addressable memory to an extended storage cache.
The buffer pools perform the first-level caching, and any extended storage cache is
used by the buffer pools as secondary-level caching. Ideally, the buffer pools hold
the data that is most frequently accessed, while the extended storage cache hold
data that is accessed less frequently.
Note: You can allocate Windows® 2000 Address Windowing Extensions (AWE)
buffer pools using the DB2_AWE registry variable. Windows AWE is a set of
memory management extensions that allow applications to manipulate
memory above certain limits, which depend on the process model of the
application. For information, refer to your Windows system documentation.
Note, however, that if you use the memory for this purpose you cannot also
use the extended storage cache.
The following database configuration parameters influence the amount and the
size of the memory available for extended storage:
v num_estore_segs defines the number of extended storage memory segments. The
default for this configuration parameter is zero, which specifies that no extended
storage cache exists.
v estore_seg_sz defines the size of each extended memory segment. This size is
determined by the platform on which the extended storage cache is used.
Note: If you use buffer pools defined with different page sizes, any of these buffer
pools can be defined to use extended storage. The page size used with
extended storage support is the largest of those defined.
Although the database manager cannot directly manipulate data that resides in the
extended storage cache, it can transfer data from the extended storage cache to the
buffer pool much faster than from disk storage.
When a row of data is needed from a page in an extended storage cache, the entire
page is read into the corresponding buffer pool.
A buffer pool and its defined associated extended storage cache are allocated when
a database is activated or when the first connection occurs.
Related concepts:
v “Buffer pool management” on page 220
v “Memory management” on page 32
Page-cleaner agents
If more pages have been written to disk, recovery of the database is faster after a
system crash because the database manager can rebuild more of the buffer pool
from disk instead of having to replay transactions from the database log files.
The size of the log that must be read during recovery is the difference between the
location of the following records in the log:
v The most recently written log record
v The log record that describes the oldest change to data in the buffer pool.
The default behavior of the page cleaners is that page cleaning is performed if the
size of the log that would need to be replayed during recovery exceeds the
following maximum:
logfilsiz * softmax
where:
v logfilsiz represents the size of the log files
| To minimize log read time during recovery, use the database system monitor to
| track the number of times that page cleaning is performed. The system monitor
| pool_lsn_gap_clns (buffer pool log space cleaners triggered) monitor element provides
| this information if you have not enabled proactive page cleaning for your database.
| If you have enabled this alternate page cleaning, this condition should not occur
| and the pool_lsn_gap_clns monitor element is always 0.
Related concepts:
v “Illustration of buffer pool data-page management” on page 225
Related reference:
v “Performance variables” on page 506
Related concepts:
v “Buffer pool management of data pages” on page 223
Database Agent
Database Agent
Buffer Pool 3. Now I can
put this page in
A A
Database Agent
Buffer Pool
There is room for
this page
Take out
A A dirty pages Write the
pages to disk
Database Agent
Asynchronous
Page Cleaner
Figure 23. Asynchronous page cleaner. “Dirty” pages are written out to disk.
Related concepts:
v “Buffer pool management” on page 220
v “Buffer pool management of data pages” on page 223
A new database has a default buffer pool called IBMDEFAULTBP with a size
determined by the platform and a default page size of 4 KB. When you create a
table space with a page size of 4 KB and do not assign it to a specific buffer pool,
Note: During normal database manager operation, you can use the ALTER
BUFFERPOOL command to resize a buffer pool.
After you create or migrate a database, you can create other buffer pools. For
example, when planning your database, you might have determined that 8 KB
page sizes were best for tables. As a result, you should create a buffer pool with an
8 KB page size as well as one or more table spaces with the same page size. You
cannot use the ALTER TABLESPACE statement to assign a table space to a buffer
pool that uses a different page size.
Note: If you create a table space with a page size greater than 4 KB, such as 8 KB,
16 KB, or 32 KB, you need to assign it to a buffer pool that uses the same
page size. If this buffer pool is currently not active, DB2® attempts to assign
the table space temporarily to another active buffer pool that uses the same
page size if one or to one of the default “hidden” buffer pools that DB2
creates when the first client connects to the database. When the database is
activated again, and the originally specified buffer pool is active, then DB2
assigns the table space to that buffer pool.
When you create a buffer pool, you specify the size of the buffer pool as a required
parameter of the DDL statement CREATE BUFFERPOOL. To increase or decrease
the buffer-pool size later, use the DDL statement ALTER BUFFERPOOL.
In a partitioned database environment, each buffer pool for a database has the
same default definition on all database partitions unless it was otherwise specified
in the CREATE BUFFERPOOL statement, or the buffer-pool size was changed by
the ALTER BUFFERPOOL statement for a particular database partition.
If any of the following conditions apply to your system, you should use only a
single buffer pool:
v The total buffer space is less than 10 000 4 KB pages.
v People with the application knowledge to do specialized tuning are not
available.
In all other circumstances, consider using more than one buffer pool for the
following reasons:
v Temporary table spaces can be assigned to a separate buffer pool to provide
better performance for queries that require temporary storage, especially
sort-intensive queries.
v If data must be accessed repeatedly and quickly by many short
update-transaction applications, consider assigning the table space that contains
the data to a separate buffer pool. If this buffer pool is sized appropriately, its
pages have a better chance of being found, contributing to a lower response time
and a lower transaction cost.
v You can isolate data into separate buffer pools to favor certain applications, data,
and indexes. For example, you might want to put tables and indexes that are
updated frequently into a buffer pool that is separate from those tables and
indexes that are frequently queried but infrequently updated. This change will
reduce the impact that frequent updates on the first set of tables have on
frequent queries on the second set of tables.
v You can use smaller buffer pools for the data accessed by applications that are
seldom used, especially for an application that requires very random access into
a very large table. In such a case, data need not be kept in the buffer pool for
longer than a single query. It is better to keep a small buffer pool for this data,
and free the extra memory for other uses, such as for other buffer pools.
v After separating different activities and data into separate buffer pools, good and
relatively inexpensive performance diagnosis data can be produced from
statistics and accounting traces.
When you use the CREATE BUFFERPOOL command to create a buffer pool or use
the ALTER BUFFERPOOL command to alter buffer pools, the total memory that is
required by all buffer pools must be available to the database manager so that all
of the buffer pools can be allocated when the database is started. If you create or
modify buffer pools while the database manager is on-line, additional memory
should be available in database global memory. If you specify the IMMEDIATE
keyword when you create a new buffer pool or increase the size of an existing
buffer pool and the required memory is not available, the database manager makes
the change the next time the database is activated. On 32-bit platforms, the
memory must be available and can be reserved in the global database memory, as
described in detailed information for the database_memory database configuration
parameter.
If this memory is not available when a database starts, the database manager
attempts to start one of each buffer pool defined with a different page size.
However, the buffer pools are started only with a minimal size of 16 pages each.
To specify a different minimal buffer-pool size, use the DB2_OVERRIDE_BPF
registry variable . Whenever a buffer pool cannot be allocated at startup, an
SQL1478W (SQLSTATE 01626) warning is returned. The database continues in this
operational state until its configuration is changed and the database can be fully
restarted.
The database manager starts with minimal-sized values only to allow you to
connect to the database so that you can reconfigure the buffer pool sizes or
perform other critical tasks. As soon as you perform these tasks, restart the
database. Do not operate the database for an extended time in such a state.
Prefetching concepts
Prefetching data into the buffer pools usually improves performance by reducing
the number of disk accesses and retaining frequently accessed data in memory.
These two methods of reading data pages are in addition to a normal read. A
normal read is used when only one or a few consecutive pages are retrieved.
During a normal read, one page of data is transferred.
The cost of inadequate prefetching is higher for parallel scans than serial scans. If
prefetching does not occur for a serial scan, the query runs more slowly because
the agent always needs to wait for I/O. If prefetching does not occur for a parallel
scan, all subagents might need to wait because one subagent is waiting for I/O.
Related concepts:
v “Buffer pool management” on page 220
v “Sequential prefetching” on page 230
v “List prefetching” on page 232
v “I/O server configuration for prefetching and parallelism” on page 233
v “Illustration of prefetching with parallel I/O” on page 234
Prefetching starts when the database manager determines that sequential I/O is
appropriate and that prefetching might improve performance. In cases such as
table scans and table sorts, the database manager can easily determine that
sequential prefetch will improve I/O performance. In these cases, the database
manager automatically starts sequential prefetch. The following example, which
probably requires a table scan, would be a good candidate for sequential prefetch:
SELECT NAME FROM EMPLOYEE
To define the number of prefetched pages for each table space, use the
PREFETCHSIZE clause in either the CREATE TABLESPACE or ALTER
TABLESPACE statements. The value that you specify is maintained in the
PREFETCHSIZE column of the [Link] system catalog table.
The database manager monitors buffer-pool usage to ensure that prefetching does
not remove pages from the buffer pool if another unit of work needs them. To
avoid problems, the database manager can limit the number of prefetched pages to
less than you specify for the table space.
The prefetch size can have significant performance implications, particularly for
large table scans. Use the database system monitor and other system monitor tools
to help you tune PREFETCHSIZE for your table spaces. You might gather
information about whether:
v There are I/O waits for your query, using monitoring tools available for your
operating system.
v Prefetch is occurring, by looking at the pool_async_data_reads (buffer pool
asynchronous data reads) data element provided by the database system monitor.
If there are I/O waits and the query is prefetching data, you might increase the
value of PREFETCHSIZE. If the prefetcher is not the cause of the I/O wait,
increasing the PREFETCHSIZE value will not improve the performance of your
query.
In some cases it is not immediately obvious that sequential prefetch will improve
performance. In these cases, the database manager can monitor I/O and activate
prefetching if sequential page reading is occurring. In this case, prefetching is
activated and deactivated by the database manager as appropriate. This type of
sequential prefetch is known as sequential detection and applies to both index and
data pages. Use the seqdetect configuration parameter to control whether the
database manager performs sequential detection.
For example, if sequential detection is turned on, the following SQL statement
might benefit from sequential prefetch:
SELECT NAME FROM EMPLOYEE
WHERE EMPNO BETWEEN 100 AND 3000
In this example, the optimizer might have started to scan the table using an index
on the EMPNO column. If the table is highly clustered with respect to this index,
then the data-page reads will be almost sequential and prefetching might improve
performance, so data-page prefetch will occur.
Index-page prefetch might also occur in this example. If many index pages must be
examined and the database manager detects that sequential page reading of the
index pages is occurring, then index-page prefetching occurs.
Related concepts:
v “Buffer pool management” on page 220
v “Prefetching data into the buffer pool” on page 229
v “List prefetching” on page 232
v “Block-based buffer pools for improved sequential prefetching” on page 231
By default, the buffer pools are page-based, which means that contiguous pages on
disk are prefetched into non-contiguous pages in memory. Sequential prefetching
can be enhanced if contiguous pages can be read from disk into contiguous pages
within a buffer pool.
You can create block-based buffer pools for this purpose. A block-based buffer pool
consist of both a page area and a block area. The page area is required for
non-sequential prefetching workloads. The block area consist of blocks where each
block contains a specified number of contiguous pages, which is referred to as the
block size.
The optimal usage of a block-based buffer pool depends on the specified block
size. The block size is the granularity at which I/O servers doing sequential
prefetching consider doing block-based I/O. The extent is the granularity at which
table spaces are striped across containers. Because multiple table spaces with
different extent sizes can be bound to a buffer pool defined with the same block
The I/O server allows some wasted pages in each buffer-pool block, but if too
much of a block would be wasted, the I/O server does non-block-based
prefetching into the page area of the buffer pool. This is not optimal performance.
For optimal performance, bind table spaces of the same extent size to a buffer pool
with a block size that equals the table-space extent size. Good performance can be
achieved if the extent size is larger than the block size, but not when the extent
size is smaller than the block size.
To create block-based buffer pools, use the CREATE and ALTER BUFFERPOOL
statements. Block-based buffer pools have the following limitations:
v A buffer pool cannot be made block-based and use extended storage
simultaneously.
v Block-based I/O and AWE support cannot be used by a buffer pool
simultaneously. AWE support takes precedence over block-based I/O support
when both are enabled for a given buffer pool. In this situation, the block-based
I/O support is disabled for the buffer pool. It is re-enabled when the AWE
support is disabled.
Note: Block-based buffer pools are intended for sequential prefetching. If your
applications do not use sequential prefetching, then the block area of the
buffer pool is wasted.
Related concepts:
v “Buffer pool management” on page 220
v “Prefetching data into the buffer pool” on page 229
v “Sequential prefetching” on page 230
List prefetching
List prefetch, or list sequential prefetch, is a way to access data pages efficiently even
when the data pages needed are not contiguous. List prefetch can be used in
conjunction with either single or multiple index access.
If the optimizer uses an index to access rows, it can defer reading the data pages
until all the row identifiers (RIDs) have been obtained from the index. For
example, the optimizer could perform an index scan to determine the rows and
data pages to retrieve, given the previously defined index IX1:
INDEX IX1: NAME ASC,
DEPT ASC,
MGR DESC,
SALARY DESC,
YEARS ASC
Related concepts:
v “Buffer pool management” on page 220
v “Prefetching data into the buffer pool” on page 229
v “Sequential prefetching” on page 230
I/O management
This section describes how to tune I/O servers for the best performance.
To estimate the number of I/O servers that you might need, consider the
following:
v The number of database agents that could be writing prefetch requests to the
I/O server queue concurrently.
v The highest degree to which the I/O servers can work in parallel.
For example, on AIX®, you might tune AIO on the operating system. When AIO
works on either SMS or DMS file containers, operating system processes called
AIO servers manage the I/O. A small number of such servers might restrict the
benefit of AIO by limiting the number of AIO requests. To configure the number of
AIO servers on AIX, use the smit AIO minservers and maxservers parameters.
Related concepts:
v “Parallel processing for applications” on page 88
Related reference:
v “num_ioservers - Number of I/O servers” on page 375
3
I/O Server Buffer Pool
Queue
1 The user application passes the SQL request to the database agent that has
been assigned to the user application by the database manager.
2, 3
The database agent determines that prefetching should be used to obtain
the data required to satisfy the SQL request and writes a prefetch request
to the I/O server queue.
Related concepts:
v “Prefetching data into the buffer pool” on page 229
v “Sequential prefetching” on page 230
v “List prefetching” on page 232
v “I/O server configuration for prefetching and parallelism” on page 233
v “Parallel I/O management” on page 235
v “Agents in a partitioned database” on page 261
Although a separate I/O server can handle the workload for each container, the
actual number of I/O servers that can perform I/O in parallel is limited to the
number of physical devices over which the requested data is spread. For this
reason, you need as many I/O servers as physical devices.
Related concepts:
v “I/O server configuration for prefetching and parallelism” on page 233
v “Illustration of prefetching with parallel I/O” on page 234
v “Guidelines for sort performance” on page 236
In general, overall sort memory available across the instance (sheapthres) should be
as large as possible without causing excessive paging. Although a sort can be
performed entirely in sort memory, this might cause excessive page swapping. In
this case, you lose the advantage of a large sort heap. For this reason, you should
use an operating system monitor to track changes in system paging whenever you
adjust the sorting configuration parameters.
Also note that in a piped sort, the sort heap is not freed until the application closes
the cursor associated with that sort. A piped sort can continue to use up memory
until the cursor is closed.
Note: With the improvement in the DB2® partial-key binary sorting technique to
include non-integer data type keys, some additional memory is required
when sorting long keys. If long keys are used for sorts, increase the sortheap
configuration parameter.
Note: You can search through the explain tables to identify the queries that have
sort operations.
You can use the database system monitor and benchmarking techniques to help set
the sortheap and sheapthres configuration parameters. For each database manager
and its databases:
v Set up and run a representative workload.
v For each applicable database, collect average values for the following
performance variables over the benchmark workload period:
– Total sort heap in use
– Active sorts
v Set sortheap to the average total sort heap in use for each database.
v Set the sheapthres. To estimate an appropriate size:
1. Determine which database in the instance has the largest sortheap value.
2. Determine the average size of the sort heap for this database.
If this is too difficult to determine, use 80% of the maximum sort heap
3. Set sheapthres to the average number of active sorts times the average size of
the sort heap computed above.
This is a recommended initial setting. You can then use benchmark
techniques to refine this value.
Related reference:
v “sortheap - Sort heap size” on page 355
v “sheapthres - Sort heap threshold” on page 354
v “sheapthres_shr - Sort heap threshold for shared sorts” on page 344
Table management
This section describes methods of managing tables for performance improvements.
Table reorganization
After many changes to table data, logically sequential data may be on
non-sequential physical data pages so that the database manager must perform
additional read operations to access data. Additional read operations are also
required if a significant number of rows have been deleted. In such a case, you
might consider reorganizing the table to match the index and to reclaim space. You
can reorganize the system catalog tables as well as database tables.
Note: Because reorganizing a table usually takes more time than running statistics,
you might execute RUNSTATS to refresh the current statistics for your data
and rebind your applications. If refreshed statistics do not improve
Consider the following factors, which might indicate that you should reorganize a
table:
v A high volume of insert, update, and delete activity on tables accessed by
queries
v Significant changes in the performance of queries that use an index with a high
cluster ratio
v Executing RUNSTATS to refresh statistical information does not improve
performance
v The REORGCHK command indicates a need to reorganize your table
v The tradeoff between the cost of increasing degradation of query performance
and the cost of reorganizing your table, which includes the CPU time, the
elapsed time, and the reduced concurrency resulting from the REORG utility
locking the table until the reorganization is complete.
To reduce the need for reorganizing a table, perform these tasks after you create
the table:
v Alter table to add PCTFREE
v Create clustering index with PCTFREE on index
v Sort the data
v Load the data
After you have performed these tasks, the table with its clustering index and the
setting of PCTFREE on table helps preserve the original sorted order. If enough
space is allowed in table pages, new data can be inserted on the correct pages to
maintain the clustering characteristics of the index. As more data is inserted and
the pages of the table become full, records are appended to the end of the table so
that the table gradually becomes unclustered.
If you perform a REORG TABLE or a sort and LOAD after you create a clustering
index, the index attempts to maintain a particular order of data, which improves
the CLUSTERRATIO or CLUSTERFACTOR statistics collected by the RUNSTATS
utility.
Note: Creating multi-dimensional clustering (MDC) tables might reduce the need
to reorganize tables. For MDC tables, clustering is maintained on the
columns that you specify as arguments to the ORGANIZE BY DIMENSIONS
clause of the CREATE TABLE statement. However, REORGCHK might
recommend reorganization of an MDC table if it considers that there are too
many unused blocks or that blocks should be compacted.
Related concepts:
v “Index reorganization” on page 252
v “DMS device considerations” on page 255
v “SMS table spaces” on page 14
v “DMS table spaces” on page 15
v “Table and index management for standard tables” on page 18
v “Snapshot monitor” in the System Monitor Guide and Reference
Related tasks:
v “Determining when to reorganize tables” on page 240
v “Choosing a table reorganization method” on page 242
Note: The REORGCHK command also returns statistical information about data
organization and can advise you about whether particular tables need to be
reorganized. However, running specific queries against the catalog statistics
tables at regular intervals or specific times can provide a performance
history that allows you to spot trends that might have wider implications
for performance.
Procedure:
To determine whether you need to reorganize tables, query the catalog statistics
tables and monitor the following statistics:
1. Overflow of rows
Query the OVERFLOW column in the [Link] table to monitor the
overflow value. The values in this column represent the number of rows that
do not fit on their original pages. Row data can overflow when VARCHAR
columns are updated with values that are longer than the initial values. In such
cases, a pointer is kept at the original location in the row and the actual value
is stored in another location that is indicated by the pointer. This can impact
performance because the database manager must follow the pointer to find the
contents of the row. This two-step process increases the processing time and
might also increase the number of I/Os required.
Reorganizing the table data will eliminate the row overflows; therefore, as the
number of overflow rows increases, the potential benefit of reorganizing your
table data increases.
2. Fetch statistics
Query the three following columns in the [Link] and
[Link] catalog statistics tables to determine the effectiveness of the
Note: In general, only one of the indexes in a table can have a high degree of
clustering.
Index scans that are not index-only accesses might perform better with higher
cluster ratios. A low cluster ratio leads to more I/O for this type of scan, since
Note: By default, ten percent free space is left on each index page when the
indexes are built. To increase the free space amount, specify the
PCTFREE parameter when you create the index. Whenever you
reorganize the index, the PCTFREE value is used. Free space greater than
ten percent might reduce frequency of index reorganization because the
additional space can accommodate additional index inserts.
6. Comparison of file pages
To calculate the number of empty pages in a table, query the FPAGES and
NPAGES columns in [Link] and subtract the NPAGES number from
the FPAGES number. The FPAGES column stores the total number of pages in
use; the NPAGES column stores the number of pages that contain rows. Empty
pages can occur when entire ranges of rows are deleted.
As the number of empty pages increases, the need for a table reorganization
increases. Reorganizing the table reclaims the empty pages and reduces the
amount of space used by a table. In addition, because empty pages are read
into the buffer pool for a table scan, reclaiming unused pages can improve the
performance of a table scan.
Related concepts:
v “Catalog statistics tables” on page 106
v “Table reorganization” on page 238
v “Index reorganization” on page 252
Related tasks:
v “Collecting catalog statistics” on page 98
Procedure:
Note: In-place table reorganization is allowed only on tables with type-2 indexes
and without extended indexes.
Consider the following trade-offs:
– Imperfect index reorganization
You might need to reorganize indexes later to reduce index fragmentation and
reclaim index object space.
– Longer time to complete
When required, in-place reorganization defers to concurrent applications. This
means that long-running statements or RR and RS readers in long-running
applications can slow the reorganization progress. In-place reorganization
might be faster in an OLTP environment in which many small transactions
occur.
– Requires more log space
Refer to the REORG TABLE syntax descriptions for detailed information about
executing these table reorganization methods.
You can also use table snapshots to monitor the progress of table reorganization.
Table reorganization monitoring data is recorded regardless of the Database
Monitor Table Switch setting.
If an error occurs, an SQLCA dump is written to the history file. For an in-place
table reorganization, the status is recorded as PAUSED.
Related concepts:
v “Table reorganization” on page 238
v “Index reorganization” on page 252
Related tasks:
v “Determining when to reorganize tables” on page 240
Index management
The following sections describe index reorganization for performance
improvements.
You must also execute the RUNSTATS utility to collect new statistics about the
indexes in the following circumstances:
v After you create an index
v After you change the prefetch size
Note: To determine whether an index is used in a specific package, use the SQL
Explain facility. To plan indexes, use the Design Advisor from the Control
Center or the db2advis tool to get advice about indexes that might be used
by one or more SQL statements.
If no index exists on a table, a table scan must be performed for each table
referenced in a database query. The larger the table, the longer a table scan takes
because a table scan requires each table row to be accessed sequentially. Although
a table scan might be more efficient for a complex query that requires most of the
rows in a table, for a query that returns only some table rows an index scan can
access table rows more efficiently.
The optimizer chooses an index scan if the index columns are referenced in the
SELECT statement and if the optimizer estimates that an index scan will be faster
than a table scan. Index files generally are smaller and require less time to read
than an entire table, particularly as tables grow larger. In addition, the entire index
may not need to be scanned. The predicates that are applied to the index reduce
the number of rows to be read from the data pages.
Each index entry contains a search-key value and a pointer to the row containing
that value. If you specify the ALLOW REVERSE SCANS parameter in the CREATE
INDEX statement, the values can be searched in both ascending and descending
order. It is therefore possible to bracket the search, given the right predicate. An
index can also be used to obtain rows in an ordered sequence, eliminating the need
for the database manager to sort the rows after they are read from the table.
In addition to the search-key value and row pointer, an index can contain include
columns, which are non-indexed columns in the indexed row. Such columns might
make it possible for the optimizer to get required information only from the index,
without accessing the table itself.
Note: The existence of an index on the table being queried does not guarantee an
ordered result set. Only an ORDER BY clause ensures the order of a result
set.
Although indexes can reduce access time significantly, they can also have adverse
effects on performance. Before you create indexes, consider the effects of multiple
indexes on disk space and processing time:
v Each index requires storage or disk space. The exact amount depends on the size
of the table and the size and number of columns in the index.
v Each INSERT or DELETE operation performed on a table requires additional
updating of each index on that table. This is also true for each UPDATE
operation that changes the value of an index key.
v The LOAD utility rebuilds or appends to any existing indexes.
The indexfreespace MODIFIED BY parameter can be specified on the LOAD
command to override the index PCTFREE used when the index was created.
Related concepts:
v “Space requirements for indexes” in the Administration Guide: Planning
v “Index planning tips” on page 246
v “Index performance tips” on page 248
v “The Design Advisor” on page 201
v “Table reorganization” on page 238
v “Index reorganization” on page 252
v “Table and index management for standard tables” on page 18
v “Table and index management for MDC tables” on page 21
v “Index cleanup and maintenance” on page 251
Related tasks:
v “Creating an index” in the Administration Guide: Implementation
v “Collecting catalog statistics” on page 98
v “Collecting index statistics” on page 100
Use the Design Advisor from the Control Center or the db2advis tool to find the
best indexes for a specific query or for the set of queries that defines a workload.
This tool recommends indexes with such performance enhancing features as
INCLUDE columns, inherited unique indexes, and ALLOW REVERSE SCANS
indexes.
The following guidelines can help you determine how to create useful indexes for
various purposes:
v To avoid some sorts, define primary keys and unique keys, wherever possible,
by using the CREATE UNIQUE INDEX statement.
v To improve data-retrieval, add INCLUDE columns to unique indexes. Good
candidates are columns that:
– Are accessed frequently and therefore would benefit from index-only access
– Are not required to limit the range of index scans
– Do not affect the ordering or uniqueness of the index key.
v To access small tables efficiently, use indexes to optimize frequent queries to
tables with more than a few data pages, as recorded in the NPAGES column in
the [Link] catalog view. You should:
– Create an index on any column you will use when joining tables.
– Create an index on any column from which you will be searching for
particular values on a regular basis.
v To search efficiently, decide between ascending and descending ordering of keys
depending on the order that will be used most often. Although the values can be
searched in reverse direction if you specify the ALLOW REVERSE SCANS
parameter in the CREATE INDEX statement, scans in the specified index order
perform slightly better than reverse scans.
Related concepts:
v “Advantages and disadvantages of indexes” on page 244
v “Index performance tips” on page 248
v “The Design Advisor” on page 201
v “Index reorganization” on page 252
v “Online index defragmentation” on page 254
v “Multidimensional clustering (MDC) table creation, placement, and use” in the
Administration Guide: Planning
Note: Clustering is not currently maintained during updates unless you are
using range-clustered tables. That is, if you update a record so that its key
value changes in the clustering index, the record is not necessarily moved
to a new page to maintain the clustering order. To maintain clustering,
use DELETE and then INSERT instead of UPDATE.
v Keep table and index statistics up-to-date
After you create a new index, run the RUNSTATS utility to collect index
statistics. These statistics allow the optimizer to determine whether using the
index can improve access performance.
v Enable online index defragmentation
Online index defragmentation is enabled if the MINPCTUSED clause is set to
greater than zero for the index. Online index defragmentation allows indexes to
Note: The PCTFREE specified when you create the index is retained when the
index is reorganized.
Dropping and re-creating or reorganizing the index also creates a new set of
pages that are roughly contiguous and sequential and improves index page
prefetch. Although more costly in time and resources, the REORG TABLE utility
also ensures clustering of the data pages. Clustering has greater benefit for index
scans that access a significant number of data pages.
In a symmetric multi-processor (SMP) environment, if the intra_parallel database
manager configuration parameter is YES or ANY, the “classic” REORG TABLE
mode, which uses a shadow table for fast table reorganization, can use multiple
processors to rebuild the indexes.
v Analyze EXPLAIN information about index usage
Periodically, run EXPLAIN on your most frequently used queries and verify that
each of your indexes is used at least once. If an index is not used in any query,
consider dropping that index.
EXPLAIN information also lets you see if table scans on large tables are
processed as the inner table of nested loop joins. If they are, an index on the
join-predicate column is either missing or considered ineffective for applying the
join predicate.
v Use volatile tables for tables that vary widely in size
A volatile table is a table that might vary in size at run time from empty to very
large. For this kind of table, in which the cardinality varies greatly, the optimizer
might generate an access plan that favors a table scan instead of an index scan.
Declaring a table “volatile” using the ALTER TABLE...VOLATILE statement
allows the optimizer to use an index scan on the volatile table. The optimizer
will use an index scan instead of a table scan regardless of the statistics in the
following circumstances:
– All columns referenced are in the index
– The index can apply a predicate in the index scan.
If the table is a typed table, using the ALTER TABLE...VOLATILE statement is
supported only on the root table of the typed table hierarchy.
Related concepts:
v “Advantages and disadvantages of indexes” on page 244
v “Index planning tips” on page 246
v “Index structure” on page 23
v “Index access and cluster ratios” on page 153
Related reference:
v “intra_parallel - Enable intra-partition parallelism” on page 449
Note: In DB2® Version 8.1 and later, all new indexes are created as type-2 indexes.
The one exception is when you add an index on a table that already has
type-1 indexes. In this case only, the new index will also be a type-1 index.
To find out what type of index exists for a table, execute the INSPECT
command. To convert type-1 indexes to type-2 indexes, execute the REORG
INDEXES command.
Index keys that are marked deleted are cleaned up in the following circumstances:
v During subsequent insert, update, or delete activity
During key insertion, keys that are marked deleted and are known to be
committed are cleaned up if such a cleanup might avoid the need to perform a
page split and prevent the index from increasing in size.
Related concepts:
v “Advantages and disadvantages of indexes” on page 244
v “Index planning tips” on page 246
v “Index structure” on page 23
v “Table reorganization” on page 238
v “Index reorganization” on page 252
Index reorganization
As tables are updated with deletes and inserts, index performance degrades in the
following ways:
v Fragmentation of leaf pages
When leaf pages are fragmented, I/O costs increase because more leaf pages
must be read to fetch table pages.
v The physical index page order no longer matches the sequence of keys on those
pages, which is referred to as a badly clustered index.
When leaf pages are badly clustered, sequential prefetching is inefficient and
results in more I/O waits.
v The index develops more than its maximally efficient number of levels.
In this case, the index should be reorganized.
If you set the MINPCTUSED parameter when you create an index, the database
server automatically merges index leaf pages if a key is deleted and the free space
When you use the REORG INDEXES command with the ALLOW WRITE ACCESS
option, all indexes on the specified table are rebuilt while read and write access to
the table is allowed. Any changes made to the underlying table that would affect
indexes while the reorganization is in progress are logged in the DB2® logs. In
addition, the same changes are placed in the internal memory buffer space, if there
is any such memory space available for use. The reorganization will process the
logged changes to catch up with current writing activity while rebuilding the
indexes. The internal memory buffer space is a designated memory area allocated
on demand from the utility heap to store the changes to the index being created or
reorganized. The use of the memory buffer space allows the index reorganization
to process the changes by directly reading from memory first, and then reading
through the logs if necessary, but at a much later time. The allocated memory is
freed once the reorganization operation completes. Following the completion of the
reorganization, the rebuilt index might not be perfectly clustered. If PCTFREE is
specified for an index, that percent of space is preserved on each page during
reorganization.
Note: The CLEANUP ONLY option of the REORG INDEXES command does not
fully reorganize indexes. The CLEANUP ONLY ALL option removes keys
that are marked deleted and are known to be committed. It also frees pages
in which all keys are marked deleted and are known to be committed. When
pages are freed, adjacent leaf pages are merged if doing so can leave at least
PCTFREE free space on the merged page. PCTFREE is the percentage of free
space defined for the index when it is created. The CLEANUP ONLY PAGES
option deletes only pages in which all keys are marked deleted and are
known to be committed.
Note: If a REORG INDEXES ALL with the ALLOW NO ACCESS option fails,
the indexes are marked bad and the operation is not undone. However, if
a REORG with the ALLOW READ ACCESS or a REORG with the
ALLOW WRITE ACCESS option fails, the original index object is restored.
Related tasks:
v “Choosing a table reorganization method” on page 242
Pages in the index are freed when the last index key on the page is removed. The
exception to this rule occurs when you specify MINPCTUSED clause in the CREATE
INDEX statement. The MINPCTUSED clause specifies a percent of space on an index
leaf page. When an index key is deleted, if the percent of filled space on the page
is at or below the specified value, then the database manager tries to merge the
remaining keys with keys on an adjacent page. If there is sufficient space on an
adjacent page, the merge is performed and an index leaf page is deleted.
Index non-leaf pages are not merged during an online index defragmentation.
However, empty non-leaf pages are deleted and made available for re-use by other
indexes on the same table. To free these non-leaf pages for other objects in a DMS
storage model or to free disk space in an SMS storage model, perform a full
reorganization of the table or indexes. Full reorganization of the table and indexes
can make the index as small as possible. Index non-leaf pages are not merged
during an online index defragmentation, but are deleted and freed for re-use if
they become empty. The number of levels in the index and the number of leaf and
non-leaf pages might be reduced.
For type-2 indexes, keys are removed from a page during key deletion only when
there is an X lock on the table. During such an operation, online index
defragmentation will be effective. However, if there is not an X lock on the table
during key deletion, keys are marked deleted but are not physically removed from
the index page. As a result, no defragmentation is attempted.
To defragment type-2 indexes in which keys are marked deleted but remain in the
physical index page, execute the REORG INDEXES command with the CLEANUP
ONLY ALL option. The CLEANUP ONLY ALL option defragments the index,
Related concepts:
v “Advantages and disadvantages of indexes” on page 244
v “Index performance tips” on page 248
v “Index structure” on page 23
v “Index reorganization” on page 252
Agent management
This section describes how the database manager uses agents and how to manage
agents for good performance.
Database agents
For each database that an application accesses, various processes or threads start to
perform the various application tasks. These tasks include logging, communication,
and prefetching.
Database agents are engine dispatchable unit (EDU) processes or threads. Database
agents do the work in the database manager that applications request. In UNIX®
environments, these agents run as processes. In Intel-based operating systems such
Windows®, the agents run as threads.
A worker agent carries out application requests but has no permanent attachment to
any particular application. The coordinator worker agent has all the information
and control blocks required to complete actions within the database manager that
were requested by the application.
Agents that are not performing work for any applications and that are waiting to
be assigned are considered to be idle agents and reside in an agent pool. These
agents are available for requests from coordinator agents operating for client
programs or for subagents operating for existing coordinator agents. The number
of available agents depends on the database manager configuration parameters
maxagents and num_poolagents.
When an agent finishes its work but still has a connection to a database, it is
placed in the agent pool. Regardless of whether the connection concentrator is
enabled for the database, if an agent is not waked up to serve a new request
within a certain period of time and the current number of active and pooled agents
is greater than num_poolagents, the agent is terminated.
Agents from the agent pool (num_poolagents) are re-used as coordinator agents for
the following kinds of applications:
v Remote TCP/IP-based applications
v Local applications on UNIX-based operating systems
v Both local and remote applications on Windows operating systems.
Other kinds of remote applications always create a new agent. If no idle agents
exist when an agent is required, a new agent is created dynamically. Because
creating a new agent requires a certain amount of overhead CONNECT and
ATTACH performance is better if an idle agent can be activated for a client.
Related concepts:
v “Database agent management” on page 258
v “Agents in a partitioned database” on page 261
v “Connection-concentrator improvements for client connections” on page 259
The ability to control these factors separately is provided by two database manager
configuration parameters:
v The max_connections parameter, which specifies the number of connected
applications
v The max_coordagents parameter, which specifies the number of application
requests that can be processed
Because each active coordinator agents requires global resource overhead, the
greater the number of these agents the greater the chance that the upper limits of
available database global resources will be reached. To prevent reaching the upper
limits of available database global resources, you might set the value of
max_connections higher than the value of max_coordagents.
Related concepts:
v “Agents in a partitioned database” on page 261
v “Connection-concentrator improvements for client connections” on page 259
v “Configuration parameters that affect the number of agents” on page 258
Related concepts:
v “Database agents” on page 256
v “Database agent management” on page 258
v “Agents in a partitioned database” on page 261
With DB2Connect connection pooling and the connection concentrator, the active
agent does not close its outbound connection after a client disconnects, but is
placed in the agent pool for the application, where it becomes a logical subagent,
which is controlled by a logical coordinator agent. with an active connection to the
remote host.
Usage examples:
1. Consider an ESE environment with a single database partition in which, on
average, 1000 users are connected to the database. At times, the number of
concurrent transactions is as high as 200, but never higher than 250.
Transactions are short.
For this workload, the administrator sets the following database manager
configuration parameters:
v max_connections is set to 1000 to ensure support for the average number of
connections.
v max_coordagents is set to 250 to support the maximum number of concurrent
transactions.
v maxagents is set high enough to support all of the coordinator agents and
subagents (where applicable) that are required to execute transactions on the
node.
If intra_parallel is OFF, maxagents is set to 250 because in such an
environment, there are no subagents. If intra_parallel is ON, maxagents should
be set large enough to accommodate the coordinator agent and the subagents
required for each transaction that accesses data on the node. For example, if
each transaction requires 4 subagents, maxagents should be set to (4+1) * 250,
which is 1250. To tune maxagents further, take monitor snapshots for the
database manager. The high-water mark of the agents will indicate the
appropriate setting for maxagents.
v num_poolagents is set to at least 250, or as high as 1250, depending on the
value of maxagents to ensure that enough database agents are available to
service incoming client requests without the overhead of creating new ones.
However, this number could be lowered to reduce resource usage during
low-usage periods. Setting this value too low causes agents to be deallocated
instead of going into the agent pool, which requires new agents to be created
before the server is able to handle an average workload.
v num_init_agents is set to be the same as num_poolagents because you know
the number of agents that should be active. This causes the database to
create the appropriate number of agents when it starts instead of creating
them before a given request can be handled.
The ability of the underlying hardware to handle a given workload is not
discussed here. If the underlying hardware cannot handle X-number of agents
working at the same time, then you need to reduce this number to the
maximum that the hardware can support. For example, if the maximum is only
1500 agents, then this limits the maximum number of concurrent transactions
that can be handled. You should monitor this kind of performance-related
setting because it is not always possible to determine exact requests sent to
other nodes at at given point in time.
2. In a system in which the workload needs to be restricted to a maximum 100
concurrent transactions and the same number of connected users as in example
1, you can set database manager configuration parameters as follows:
v max_coordagents is set to 100
v num_poolagents is set to 100
With these settings, the maximum number of clients that can concurrently
execute transactions is 100. When all clients disconnect, 100 agents are waiting
Related concepts:
v “Database agents” on page 256
v “Database agent management” on page 258
v “DB2 architecture and process overview” on page 9
v “Memory management” on page 32
Related concepts:
v “I/O server configuration for prefetching and parallelism” on page 233
v “Illustration of prefetching with parallel I/O” on page 234
v “Database agents” on page 256
v “Database agent management” on page 258
v “Configuration parameters that affect the number of agents” on page 258
Because collecting some of this data introduces overhead on the operation of DB2,
monitor switches are available to control which information is collected. To set
You can access the data that the database manager maintains either by taking a
snapshot or by using an event monitor.
Taking a snapshot
An event monitor captures system monitor information after particular events have
occurred, such as the end of a transaction, the end of a statement, or the detection
of a deadlock. This information can be written to files or to a named pipe.
Note: If the database system that you are monitoring is not running on the
same machine as the Control Center, you must copy the event monitor
file to the same machine as the Control Center before you can view
the trace. An alternative method is to place the file in a shared file
system accessible to both machines.
Related concepts:
v “Quick-start tips for performance tuning” on page 7
A governor instance consists of a front-end utility and one or more daemons. Each
instance of the governor that you start is specific to an instance of the database
manager. By default, when you start the governor a governor daemon starts on
each partition of a partitioned database. However, you can specify that a daemon
be started on a single partition that you want to monitor.
Note: When the governor is active, its snapshot requests might affect database
manager performance. To improve performance, increase the governor
wake-up interval to reduce its CPU usage.
Each governor daemon collects information about the applications that run against
the database. If then checks this information against the rules that you specify in
the governor configuration file for this database.
If the action associated with a rule changes the priority of the application, the
governor changes the priority of agents on the database partition where the
resource violation occurred. In a partitioned database, if the application is forced to
disconnect from the database, the action occurs even if the daemon that detected
the violation is running on the coordinator node of the application.
The governor logs any actions that it takes. To review the actions, you query the
log files.
Related concepts:
v “The Governor daemon” on page 267
v “The governor configuration file” on page 269
v “Governor log files” on page 276
Related tasks:
v “Starting and stopping the governor” on page 266
v “Configuring the Governor” on page 268
Related reference:
v “db2gov - DB2 Governor Command” in the Command Reference
Prerequisites:
Before you start the governor, you must create the configuration file.
Restrictions:
To start or stop the governor, you must have sysadm or sysctrl authorization.
Procedure:
Related reference:
v “db2gov - DB2 Governor Command” in the Command Reference
Note: On some platforms, the CPU statistics are not available from the DB2®
Monitor. In this case, the account rule and the CPU limit are not
available.
3. It checks the statistics for each application against the rules in the governor
configuration file. If a rule applies to an application, the governor performs the
specified action.
Note: The governor cannot be used to adjust agent priorities if the agentpri
database manager configuration parameter is anything other than the system
default. (This note does not apply to Windows® NT platforms.)
When the governor finishes its tasks, it sleeps for the interval specified in the
configuration file. when the interval elapses, the governor wakes up and begins the
task loop again.
When the governor encounters an error or stop signal, it does cleanup processing
before it ends. Using a list of applications whose priorities have been set, the
cleanup processing resets all application agent priorities. It then resets the priorities
of any agents that are no longer working on an application. This ensures that
agents do not remain running with nondefault priorities after the governor ends. If
an error occurs, the governor writes a message to the administration notification
log to indicate that it ended abnormally.
Note: Although the governor daemon is not a database application, and therefore
does not maintain a connection to the database, it does have an instance
attachment. Because it can issue snapshot requests, the governor daemon
can detect when the database manager ends.
Related concepts:
Related tasks:
v “Starting and stopping the governor” on page 266
Governor configuration
This section explains how to configure the governor to monitor and control
database activity.
The configuration file consists of a set of rules. The first three rules specify the
database to monitor, the interval at which to write log records, and the interval at
which to wake up for monitoring. The remaining rules specify how to monitor the
database server and what actions to take in specific circumstances.
Procedure:
Related concepts:
v “The governor configuration file” on page 269
v “Governor rule elements” on page 271
Related reference:
v “db2gov - DB2 Governor Command” in the Command Reference
If your rule requirements change, you edit the configuration file without stopping
the governor. Each governor daemon detects that the file has changed, and rereads
it.
The configuration file must be created in a directory that is mounted across all the
database partitions so that the governor daemon on each partition can read the
same configuration file.
The configuration file consists of three required rules that identify the database to
be monitored, the interval at which log records are written, and the sleep interval
of the governor daemons. Following these parameters, the configuration file
contains a set of optional application-monitoring rules and actions. The following
comments apply to all rules:
v Delimit comments inside { } braces.
v Most entries can be specified in uppercase, lowercase, or mixed case characters.
The exception is the application name, specified as an argument to the applname
rule, which is case sensitive.
v Each rule ends with a semicolon (;).
Required rules
The following rules specify the database to be monitored and the interval at which
the daemon wakes up after each loop of activities. Each of these rules is specified
only once in the file.
dbname
The name or alias of the database to be monitored.
account nnn
Account records are written containing CPU usage statistics for each
connection at the specified number of minutes.
Following the required rules, you can add rules that specify how to govern the
applications. These rules are made of smaller components called rule clauses. If
used, the clauses must be entered in a specific order in the rule statement, as
follows:
1. desc (optional): a comment about the rule, enclosed in quotation marks
2. time (optional): the time during the day when the rule is evaluated
3. authid (optional): one or more authorization IDs under which the application
executes statements
4. applname (optional): the name of the executable or object file that connects to
the database. This name is case sensitive. The application name must be
surrounded by double quotes if the application contains spaces.
5. setlimit: the limits that the governor checks. These can be one of several, for
example, CPU time, number of rows returned, or idle time..
6. action (optional): the action to take if a limit is reached. If no action is
specified, the governor reduces the priority of agents working for the
application by 10 when a limit is reached. Actions against the application can
include reducing its agent priority, forcing it to disconnect from the database, or
setting scheduling options for its operations.
You combine the rule clauses to form a rule, using each clause only once in each
rule, and end the rule with a semicolon, as shown in the following examples:
desc "Allow no UOW to run for more than an hour"
setlimit uowtime 3600 action force;
desc "Slow down the use of db2 CLP by the novice user"
authid novice
applname [Link]
setlimit cpu 5 locks 100 rowssel 250;
If more than one rule applies to an application, all are applied. Usually, the action
associated with the rule limit encountered first is the action that is applied first. An
exception occurs you specify if -1 for a clause in a rule. In this case, the value
specified for the clause in the subsequent rule can only override the value
previously specified for the same clause: other clauses in the previous rule are still
operative. For example, one rule uses the rowssel 100000 uowtime 3600 clauses to
specify that the priority of an application is decreased either if its elapsed time is
greater than 1 hour or if it selects more than 100 000 rows. A subsequent rule uses
the uowtime -1 clause to specify that the same application can have unlimited
elapsed time. In this case, if the application runs for more than 1 hour, its priority
is not changed. That is, uowtime -1 overrides uowtime 3600. However, if it selects
more than 100 000 rows, its priority is lowered because rowssel 100000 is still
valid.
The governor processes rules in the configuration file from the top of the file to the
bottom. However, if a later rule’s setlimit clause is more relaxed than a preceding
rule, the more restrictive rule still applies. For example, in the following
configuration file, admin will be limited to 5000 rows despite the later rule because
the first rule is more restrictive.
To ensure that a less restrictive rule overrides a more restrictive rule that occurs
earlier in the file, you can specify the -1 option to clear the previous rule before
applying the new one. For example, in the following configuration file, the initial
rule limits all users to 5000 rows. The second rule clears this limit for admin, and
the third rule resets the limit for admin to 10000 rows.
desc "Force anyone selecting 5000 or more rows"
setlimit rowssel 5000 action force;
Related concepts:
v “Governor rule elements” on page 271
v “Example of a Governor configuration file” on page 275
Limit clauses
setlimit
Specifies one or more limits for the governor to check. The limits can only
be -1 or greater than 0 (for example, cpu -1 locks 1000 rowssel 10000). At
least one of the limits (cpu, locks, rowsread, uowtime) must be specified,
and any limit not specified by the rule is not limited by that particular
rule. The governor can check the following limits:
cpu nnn
Specifies the number of CPU seconds that can be consumed by an
application. If you specify -1, the governor does not limit the
application’s CPU usage.
Action clauses
[action]
Specifies the action to take if one or more of the specified limits is
exceeded. You can specify the following actions.
Note: If a limit is exceeded and the action clause is not specified, the
governor reduces the priority of agents working for the application
by 10.
nice nnn
Specifies a change to the priority of agents working for the
application. Valid values are from −20 to +20.
For this parameter to be effective:
v On UNIX®-based platforms, the agentpri database manager
parameter must be set to the default value; otherwise, it
overrides the priority clause.
v On Windows platforms, the agentpri database manager
parameter and priority action may be used together.
force Specifies to force the agent that is servicing the application. (Issues
a FORCE APPLICATION to terminate the coordinator agent.)
schedule [class]
Scheduling improves the priorities of the agents working on the
applications with the goal of minimizing the average response
times while maintaining fairness across all applications.
The governor chooses the top applications for scheduling based on
the following three criteria:
v The application holding the most locks
This choice is an attempt to reduce the number of lockwaits.
v The oldest application
v The application with the shortest estimated remaining running
time.
This choice is an attempt to allow as many short-lived
statements as possible to complete during the interval.
Note: If a limit is exceeded and the action clause is not specified, the
governor reduces the priority of agents working for the application.
Related concepts:
v “The governor configuration file” on page 269
v “Example of a Governor configuration file” on page 275
Related tasks:
v “Configuring the Governor” on page 268
desc "Schedule all CPU hogs in one class which will control consumption"
setlimit cpu 3600
action schedule class;
desc "Slow down the use of db2 CLP by the novice user"
authid novice
applname [Link]
setlimit cpu 5 locks 100 rowssel 250;
desc "During day hours do not let anyone run for more than 10 seconds"
time 8:30 17:00 setlimit cpu 10 action force;
Related concepts:
v “The governor configuration file” on page 269
v “Governor rule elements” on page 271
Related tasks:
v “Configuring the Governor” on page 268
Each governor daemon has a separate log file. Separate log files prevent
file-locking bottlenecks that might result when many governor daemons write to
the same file at the same time. To merge the log files together and query them, use
the db2govlg utility.
Note: The format of the Date and Time fields is yyyy-mm-dd [Link]. You can
merge the log files for each database partition by sorting on this field.
The NodeNum field indicates the number of the database partition on which the
governor is running.
The RecType field contains different values, depending on the type of log record
being written to the log. The values that can be recorded are:
v START: the governor was started
v STOP: the governor was stopped
v FORCE: an application was forced
v NICE: the priority of an application was changed
v ERROR: an error occurred
v WARNING: a warning occurred
v READCFG: the governor read the configuration file
v ACCOUNT: the application accounting statistics.
v SCHEDGRP: a change in agent priorities occurred.
Some of these values are described in more detail below.
START
The START record is written when the governor is started. It has the
following format:
Database = <database_name>
STOP The STOP record is written when the governor is stopped. It has the
following format:
Database = <database_name>
FORCE
The FORCE record is written out whenever the governor determines that
an application is to be forced as required by a rule in the governor
configuration file. The FORCE record has the following format:
<appl_name> <auth_id> <appl_id> <coord_partition> <cfg_line>
<restriction_exceeded>
where:
<coord_partition>
Specifies the number of the application’s coordinating partition.
<cfg_line>
Specifies the line number in the governor configuration file where
the rule causing the application to be forced is located.
Because standard values are written, you can query the log files for different types
of actions. The Message field provides other nonstandard information that varies
according to the value under the RecType field. For instance, a FORCE or NICE record
indicates application information in the Message field, while an ERROR record
includes an error message.
Related concepts:
v “The Governor utility” on page 265
v “Governor log file queries” on page 280
db2govlg log-file
nodenum node-num rectype record-type
There are no authorization restrictions for using this utility. This allows all users to
query whether the governor has affected their application. If you want to restrict
access to this utility, you can change the group permissions for the db2govlg file.
Related concepts:
v “The Governor utility” on page 265
v “Governor log files” on page 276
If these simple strategies do not add the capacity you need, consider the following
methods:
v Add processors.
If a single-partition configuration with a single processor is used to its maximum
capacity, you might either add processors or add partitions. The advantage of
adding processors is greater processing power. In an SMP system, processors
share memory and storage system resources. All of the processors are in one
system, so there are no additional overhead considerations such as
communication between systems and coordination of tasks between systems.
Utilities in DB2® such as load, backup, and restore can take advantage of the
additional processors. DB2 Universal Database™ supports this environment.
Note: Some operating systems, such as the Solaris Operating Environment, can
dynamically turn processors on- and off-line.
If you add processors, review and modify some database configuration
parameters that determine the number of processors used. The following
database configuration parameters determine the number of processors used and
might need to be updated:
– Default degree (dft_degree)
– Maximum degree of parallelism (max_querydegree)
– Enable intra-partition parallelism (intra_parallel)
You should also evaluate parameters that determine how applications perform
parallel processing.
In an environment where TCP/IP is used for communication, review the value
for the DB2TCPCONNMGRS registry variable.
v Add physical nodes.
If your database manager is currently partitioned, you can increase both
data-storage space and processing power by adding separate single-processor or
multiple-processor physical nodes. The memory and storage system resources on
each node are not shared with the other nodes. Although adding nodes might
result in communication and task-coordination issues, this choice provides the
advantage of balancing data and user access across more than one system. DB2
Universal Database supports this environment.
You can add nodes either while the database manager system is running or
while it is stopped. If you add nodes while the system is running, however, you
must stop and restart the system before databases migrate to the new node.
When you add a new database partition, you cannot drop or create a database that
takes advantage of the new partition until the procedure is complete, and the new
server is successfully integrated into the system.
Related concepts:
v “Partitions in a partitioned database” on page 282
If your system is stopped, you use db2start. If it is running, you can use any of
the other choices.
When you use the ADD DBPARTITIONNUM command to add a new database
partition to the system, all existing databases in the instance are expanded to the
new database partition. You can also specify which containers to use for temporary
table spaces for the databases. The containers can be:
v The same as those defined for the catalog node for each database. (This is the
default.)
v The same as those defined for another database partition.
v Not created at all. You must use the ALTER TABLESPACE statement to add
temporary table space containers to each database before the database can be
used.
You cannot use a database on the new partition to contain data until one or more
database partition groups are altered to include the new database partition.
Note: If no databases are defined in the system and you are running Enterprise
Server Edition on a UNIX®-based system, edit the [Link] file to add a
new database partition definition; do not use any of the procedures
described, as they apply only when a database exists.
Related tasks:
v “Adding a partition to a running database system” on page 283
v “Adding a partition to a stopped database system on Windows NT” on page 284
v “Dropping a database partition” on page 288
Procedure:
Note: You might have to issue the DB2START command twice for all database
partition servers to access the new [Link] file.
4. Back up all databases on the new database partition. (Optional)
5. Redistribute data to the new database partition. (Optional)
Related tasks:
v “Adding a partition to a stopped database system on Windows NT” on page 284
v “Adding a partition to a stopped database system on UNIX” on page 285
Prerequisites:
You must install the new server before you can create a partition on it.
Procedure:
Note: You might have to issue the DB2START command twice for all database
partition servers to access the new [Link] file.
6. Back up all databases on the new database partition. (Optional)
7. Redistribute data to the new database partition. (Optional)
Related concepts:
v “Partitions in a partitioned database” on page 282
v “Node-addition error recovery” on page 287
Related tasks:
v “Adding a partition to a running database system” on page 283
v “Adding a partition to a stopped database system on UNIX” on page 285
Prerequisites:
You must install the new server if it does not exist, including the following tasks:
v Making executables accessible (using shared file-system mounts or local copies)
v Synchronizing operating system files with those on existing processors
v Ensuring that the sqllib directory is accessible as a shared file system
v Ensuring that the relevant operating system parameters (such as the maximum
number of processes) are set to the appropriate values
You must also register the host name with the name server or in the hosts file in
the etc directory on all database partitions.
Procedure:
Note: You might have to issue the DB2START command twice for all database
partition servers to access the new [Link] file.
6. Back up all databases on the new database partition. (Optional)
7. Redistribute data to the new database partition. (Optional)
Related concepts:
v “Node-addition error recovery” on page 287
Related tasks:
v “Adding a partition to a running database system” on page 283
Related tasks:
v “Adding a partition to a running database system” on page 283
v “Adding a partition to a stopped database system on Windows NT” on page 284
Prerequisites:
Verify that the partition is not in use by issuing the DROP NODE VERIFY
command or the sqledrpn API.
v If you receive message SQL6034W (Node not used in any database), you can
drop the partition.
v If you receive message SQL6035W (Node in use by database), use the
REDISTRIBUTE NODEGROUP command to redistribute the data from the
database partition that you are dropping to other database partitions from the
database alias.
Also ensure that all transactions for which this database partition was the
coordinator have all committed or rolled back successfully. This may require doing
crash recovery on other servers. For example, if you drop the coordinator database
partition (that is, the coordinator node), and another database partition
participating in a transaction crashed before the coordinator node was dropped,
the crashed database partition will not be able to query the coordinator node for
the outcome of any in-doubt transactions.
Procedure:
Related concepts:
v “Management of database server capacity” on page 281
v “Partitions in a partitioned database” on page 282
Data redistribution
To redistribute table data among the partitions in a partitioned database, you use
the REDISTRIBUTE DATABASE PARTITION GROUP command.
In a partitioned database you might redistribute data for the following reasons:
v To balance data volumes and processing loads across database partitions.
Performance improves if data access can be spread out over more than one
partition.
v To introduce skew in the data distribution across database partitions.
Access and throughput performance might improve if you redistribute data in a
frequently accessed table so that infrequently accessed data is on a small number
of database partitions in the database partitioning group, and the frequently
accessed data is distributed over a larger number of partitions. This would
improve access performance and throughput on the most frequently run
applications.
Related concepts:
v “Log space requirements for data redistribution” on page 293
v “Redistribution error recovery” on page 294
Related tasks:
v “Redistributing data across partitions” on page 291
v “Determining whether to redistribute data” on page 290
Procedure:
Note: You can also use AutoLoader utility with its ANALYZE option to create a
data distribution file. You can use this file as input to the Data
Redistribution utility.
Related concepts:
v “Data redistribution” on page 289
v “Log space requirements for data redistribution” on page 293
Related tasks:
v “Redistributing data across partitions” on page 291
Prerequisites:
Log file size: Ensure that log files are large enough for the data redistribution
operation. The log file on each affected partition must be large enough to
accommodate the INSERT and DELETE operations performed there.
Restrictions:
You can do the following operations on objects of the database partition group
while the utility is running. You cannot, however, do them on the table that is
being redistributed. You can:
v Create indexes on other tables. The CREATE INDEX statement uses the
partitioning map of the affected table.
You cannot use this procedure to redistribute data after adding a partition to a
single-partition system unless all affected tables have a partitioning key. The
REDISTRIBUTE DATABASE PARTITION GROUP command relies on partitioning
keys to redistribute data. The partitioning key is generated automatically when a
table is created in a multi-partition database partition group, or can be explicitly
defined using the CREATE TABLE or ALTER TABLE SQL statements. If your tables
were created in a single-partition partition group, and you did not define the
partitioning key in the CREATE TABLE SQL statement, there will be no
partitioning keys defined. You must use the ALTER TABLE SQL statement to create
a partitioning key for each affected table before redistributing the data.
Procedure:
Note: The Explain tables contain information about the partitioning map used to
redistribute data.
Related concepts:
v “Data redistribution” on page 289
v “Log space requirements for data redistribution” on page 293
v “Redistribution error recovery” on page 294
Related tasks:
v “Determining whether to redistribute data” on page 290
The log must be large enough to accommodate the INSERT and DELETE
operations at each database partition where data is being redistributed. The
heaviest logging requirements will be either on the database partition that will lose
the most data, or on the database partition that will gain the most data.
If you are moving to a larger number of database partitions, use the ratio of
current database partitions to the new number of database partitions to estimate
the number of INSERT and DELETE operations. For example, consider
redistributing data that is uniformly distributed before redistribution. If you are
moving from four to five database partitions, approximately twenty percent of the
four original database partitions will move to the new database partition. This
means that twenty percent of the DELETE operations will occur on each of the
four original database partitions, and all of the INSERT operations will occur on
the new database partition.
Consider a non-uniform distribution of the data, such as the case in which the
partitioning key contains many NULL values. In this case, all rows that contain a
NULL value in the partitioning key move from one database partition under the
old partitioning scheme and to a different database partition under the new
partitioning scheme. As a result, the amount of log space required on those two
database partitions increases, perhaps well beyond the amount calculated by
assuming uniform distribution.
The redistribution of each table is a single transaction. For this reason, when you
estimate log space, you multiply the percentage of change, such as twenty percent,
by the size of the largest table. Consider, however, that the largest table might be
uniformly distributed but the second largest table, for example, might have one or
more inflated database partitions. In such a case, consider using the non-uniformly
distributed table instead of the largest one.
Related concepts:
v “Data redistribution” on page 289
Note: On non-UNIX platforms, only the first eight (8) bytes of the database
partition group name are used.
If the data redistribution operation fails, some tables may be redistributed, while
others are not. This occurs because data redistribution is performed a table at a
time. You have two options for recovery:
v Use the CONTINUE option to continue the operation to redistribute the
remaining tables.
v Use the ROLLBACK option to undo the redistribution and set the redistributed
tables back to their original state. The rollback operation can take about the
same amount of time as the original redistribution operation.
Before you can use either option, a previous data redistribution operation must
have failed such that the REBALANCE_PMID column in the
[Link] table is set to a non-NULL value.
If you happen to delete the status file by mistake, you can still attempt a
CONTINUE operation.
Related concepts:
v “Data redistribution” on page 289
Related tasks:
v “Redistributing data across partitions” on page 291
Note: The redistribute stored procedures and functions work only in partitioned
databases, where a partitioning key has been defined for each table.
The value “-2”can be used for stepSize and totalSteps in this procedure to indicate
that the number is unlimited.
Table 34. set_swrd_settings, input parameters
Name Data type Description
dbpgName VARCHAR(128) The database partition group name, against which the redistribute
process is to run.
overwriteSpec SMALLINT Bitwise field indentifier(s) from Table 32 on page 295 indicating the
target fields to be written or overwritten into the redistribute
settings registry.
redistMethod SMALLINT The number indicating the redistribute is to run using the
distribution file or target partitioning map.
pMapFile VARCHAR (255) The full path file name of the target partition map.
distFile VARCHAR (255) The full path file name of the data distribution file.
stepSize BIGINT The maximum number of rows that can be moved before a commit
must be called to prevent a log full situation. The number can be
moved in each redistribution step.
The value “-2” can be used in this procedure to indicate that the number is
unlimited.
Table 39. stepwise_redistribute_dbpg input parameters
Name Data type Description
inDBPGroup VARCHAR(128) The name of the target database partition group
inStartingPoint SMALLINT This parameter can be NULL. If it is not null and pointing to a
positive number, it is used to overwrite the ″nextStep″ value given
by the swrd settings registry. This can be a useful option when you
want to rerun SWRD from a particular step.
inNumSteps SMALLINT The number of steps to run. If not null and pointing to a positive
number, it is used to overwrite the ″numSteps″ value given by the
swrd settings registry. This can be a useful option when you want
to rerun SWRD with a different number of steps than what is
specified in the settings. For example, if there are five steps in a
scheduled stage, and the SWRD process failed at step 3, after
correcting the error condition, SWRD can be called to run the
remaining three steps.
db_partitions UDF
The db_partitions user-defined function parses through the [Link] file, and
returns a row for each partition found.
Usage example
The following is an example of a CLP script on AIX:
# -------------------------------------------------------------------------------------
# Set the database you wish to connect to
# -------------------------------------------------------------------------------------
dbName="SAMPLE"
# -------------------------------------------------------------------------------------
# Set the target database partition group name
# -------------------------------------------------------------------------------------
dbpgName="IBMDEFAULTGROUP"
# -------------------------------------------------------------------------------------
# Specify the table name and schema
# -------------------------------------------------------------------------------------
tbSchema="$USER"
tbName="STAFF"
# -------------------------------------------------------------------------------------
# Specify the name of the data distribution file
# -------------------------------------------------------------------------------------
distFile="$HOME/sqllib/function/$dbName.IBMDEFAULTGROUP_swrdData.dst"
export DB2INSTANCE=$USER
export DB2COMM=TCPIP
# -------------------------------------------------------------------------------------
# Invoke call statements in clp
# -------------------------------------------------------------------------------------
db2start
db2 -v "connect to $dbName"
# -------------------------------------------------------------------------------------
# Analysing the effect of adding a partition without applying the changes - a ’what if’
# hypothetical analysis
#
# - In the following case, the hypothesis is adding partition 40, 50 and 60 to the
# database partition group, and for partitions 10,20,30,40,50,60, using a respective
# target ratio of 1:2:1:2:1:2.
#
# NOTE: in this example only partitions 10, 20 and 30 actually exist in the database
# partition group
# -------------------------------------------------------------------------------------
db2 -v "call sysproc.analyze_log_space(’$dbpgName’, ’$tbSchema’, ’$tbName’, 2, ’ ’,
’A’, ’40,50,60’, ’10,20,30,40,50,60’, ’1,2,1,2,1,2’)"
# -------------------------------------------------------------------------------------
# Analysing the effect of droping a partition without applying the changes
#
# - In the following case, the hypothesis is dropping partition 30 from the database
# partition group, and redistributing the data in partitions 10 and 20 using a
# respective target ratio of 1 : 1
#
# NOTE: In this example all partitions 10, 20 and 30 should exist in the database
# -------------------------------------------------------------------------------------
# Generate a data distribution file to be used by the redistribute process
# -------------------------------------------------------------------------------------
db2 -v "call sysproc.generate_distfile(’$tbSchema’, ’$tbName’, ’$distFile’)"
# -------------------------------------------------------------------------------------
# Write a step wise redistribution plan into a registry
#
# Setting the 10th parameter to 1, may cause a currently running step wise redistribute
# stored procedure to complete the current step and stop, until this parameter is reset
# to 0, and the redistribute stored procedure is called again.
# -------------------------------------------------------------------------------------
db2 -v "call sysproc.set_swrd_settings(’$dbpgName’, 255, 0, ’ ’, ’$distFile’, 1000,
12, 2, 1, 0, ’10,20,30’, ’50,50,50’)"
# -------------------------------------------------------------------------------------
# Report the content of the step wise redistribution plan for the given database
# partition group.
# -------------------------------------------------------------------------------------
db2 -v "call sysproc.get_swrd_settings(’$dbpgName’, 255, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)"
# -------------------------------------------------------------------------------------
# Redistribute the database partition group "dbpgName" according to the redistribution
# plan stored in the registry by set_swrd_settings. It starting with step 3 and
# redistributes the data until 2 steps in the redistribution plan are completed.
# -------------------------------------------------------------------------------------
db2 -v "call sysproc.stepwise_redistribute_dbpg(’$dbpgName’, 3, 2)"
Related concepts:
v “Data redistribution” on page 289
Benchmark testing
Benchmark testing is a normal part of the application development life cycle. It is a
team effort that involves both application developers and database administrators
(DBAs), and should be performed against your application in order to determine
current performance and improve it. If the application code has been written as
efficiently as possible, additional performance gains might be realized from tuning
the database and database manager configuration parameters. You can even tune
application parameters to meet the requirements of the application better.
Benchmark tests are based on a repeatable environment so that the same test run
under the same conditions will yield results that you can legitimately compare.
Note: Started applications use memory even when they are minimized or idle.
This increases the probability that paging will skew the results of the
benchmark and violates the repeatability rule.
v The hardware and software used for benchmarking match your production
environment.
For benchmarking, you create a scenario and then applications in this scenario
several times, capturing key information during each run. Capturing key
information after each run is of primary importance in determining the changes
that might improve performance of both the application and the database.
Related concepts:
v “Benchmark preparation” on page 304
v “Benchmark test creation” on page 305
v “Benchmark test execution” on page 311
v “Benchmark test analysis example” on page 312
Benchmark preparation
Complete the logical design of the database against which the application runs
before you start performance benchmarking. Set up and populate tables, views,
and indexes. Normalize tables, bind application packages, and populate tables with
realistic data.
You should also have determined the final physical design of the database. Place
database manager objects in their final disk locations, size log files, determining
the location of work files and backup, and test backup procedures. In addition,
check packages to make sure that performance options such as row blocking are
enabled when possible.
You should have reached a point in application programming and testing phases
that will enable you to create your benchmark programs. Although the practical
limits of an application might be revealed during the benchmark testing, the
purpose of the benchmark described here is to measure performance, not to detect
defects or abends.
Make sure that you run benchmark tests with a production-size database. An
individual SQL statement should return as much data and require as much sorting
as in production. This rule ensures that the application will test representative
memory requirements.
Related concepts:
v “Benchmark testing” on page 303
v “Benchmark test creation” on page 305
To test the performance of specific SQL statements, you might include these
statements alone in the benchmark program along with the necessary CONNECT,
PREPARE, OPEN, and other statements and a timing mechanism.
Another factor to consider is the type of benchmark to use. One option is to run a
set of SQL statements repeatedly over a time interval. The ratio of the number of
statements executed and this time interval would give the throughput for the
application. Another option is simply to determine the time required to execute
individual SQL statements.
For all benchmark testing, you need an efficient timing system to calculate the
elapsed time, whether for individual SQL statements or the application as a whole.
To simulate applications in which individual SQL statements are executed in
isolation, it might be important to track times for CONNECT, PREPARE, and
COMMIT statements. However, for programs that process many different
statements, perhaps only a single CONNECT or COMMIT is necessary, and
focusing on just the execution time for an individual statement might be the
priority.
Although the elapsed time for each query is an important factor in performance
analysis, it might not necessarily reveal bottlenecks. For example, information on
CPU usage, locking, and buffer pool I/O might show that the application is I/O
bound and is not using the CPU to its full capacity. A benchmark program should
allow you to obtain this kind of data for a more detailed analysis if needed.
Not all applications send the entire set of rows retrieved from a query to some
output device. For example, the whole answer set might be input for another
program, so that none of the rows from the first application are sent as output.
Formatting data for screen output usually has high CPU cost and might not reflect
user need. To provide an accurate simulation, a benchmark program should reflect
the row handling of the specific application. If rows are sent to an output device,
inefficient formatting could consume the majority of CPU processing time and
misrepresent the actual performance of the SQL statement itself.
This benchmarking tool also has a CLI option. With this option, you can specify a
cache size. In the following example, db2batch is run in CLI mode with a cache
size of 30 statements:
db2batch -d sample -f [Link] -cli 30
or the
-o <options>
and
-p <perf_detail>
and
-p <perf_detail>
control option values are supported and are valid for all DB2® Universal Database
platforms.
Statement number: 1
The above sample output includes specific data elements returned by the database
system monitor.
In the next example (on UNIX®), just the materialized query table is produced.
db2batch -d sample -f [Link] -r /dev/null,
Produces just the materialized query table. Using the -r option, outfile1 was
replaced by /dev/null and outfile2 (which contains just the materialized query
table) is empty, so db2batch sends the output to the screen:
Summary of Results
==================
Elapsed Agent CPU Rows Rows
Statement # Time (s) Time (s) Fetched Printed
1 0.074 0.020 5 5
2 0.037 Not Collected 8 5
Arith. mean 0.055
Geom. mean 0.052
Figure 30. Sample Output from db2batch -- Materialized Query Table Only
Related concepts:
v “Benchmark test creation” on page 305
v “Benchmark test execution” on page 311
v “Benchmark test analysis example” on page 312
When you run the benchmark, the first iteration, which is called a warm-up run,
should be considered a separate case from the subsequent iterations, which are
called normal runs. Because the warm-up run includes some start-up activities,
such as initializing the buffer pool, and consequently, takes somewhat longer than
normal runs. Although the information from the warm-up run might be
realistically valid, it is not statistically valid. When you calculate the average
timing or CPU for a specific set of parameter values, use only the results from
normal runs.
You might consider using the Configuration Advisor to create the warm-up run of
the benchmark. The questions that the Configuration Advisor asks can provide
insight into some things to consider when you adjust the configuration of your
environment for the normal runs during your benchmark activity. You can start the
Configuration Advisor from the Control Center or by executing the db2
autoconfigure command with appropriate options.
If benchmarking uses individual queries, ensure that you minimize the potential
effects of previous queries by flushing the buffer pool. To flush the buffer pool,
read a number of pages that irrelevant to your query and to fill the buffer pool.
After you complete the iterations for a single set of parameter values, you can
change a single parameter. However, between each iteration, perform the following
tasks to restore the benchmark environment to its original state:
v . If the catalog statistics were updated for the test, make sure that the same
values for the statistics are used for every iteration.
v The data used in the tests must be consistent if it is updated by the tests. This
can be done by:
– Using the RESTORE utility to restore the entire database. The backup copy of
the database contains its previous state, ready for the next test.
– Using the IMPORT or LOAD utility to restore an exported copy of the data.
This method allows you to restore only the data that has been affected.
REORG and RUNSTATS utilities should be run against the tables and indexes
that contain this data.
v To return the application to its original state, re-bind it to the database.
You can write a driver program to help you with your benchmark testing. This
driver program could be written using a language such as REXX or, for
UNIX®-based platforms, using shell scripts.
This driver program would execute the benchmark program, pass it the
appropriate parameters, drive the test through multiple iterations, restore the
environment to a consistent state, set up the next test with new parameter values,
and collect/consolidate the test results. These driver programs can be flexible
enough that they could be used to run the entire set of benchmark tests, analyze
the results, and provide a report of the final and best parameter values for the
given test.
Related concepts:
v “Benchmark testing” on page 303
v “Benchmark preparation” on page 304
v “Examples of db2batch tests” on page 307
Note: The data in the above report is shown for illustration purposes only. It does
not represent measured results.
Analysis shows that the CONNECT (statement 01) took 1.34 seconds, the OPEN
CURSOR (statement 10) took 2 minutes and 8.15 seconds, the FETCHES (statement
15) returned seven rows with the longest delay being .28 seconds, the CLOSE
CURSOR (statement 20) took .84 seconds, and the CONNECT RESET (statement
99) took .03 seconds.
If your program can output data in a delimited ASCII format, it could later be
imported into a database table or a spreadsheet for further statistical analysis.
Note: The data in the above report is shown for illustration purposes only. It does
not represent any measured results.
Related concepts:
v “Benchmark testing” on page 303
v “Benchmark test creation” on page 305
v “Benchmark test execution” on page 311
Configuration files contain parameters that define values such as the resources
allocated to the DB2 UDB products and to individual databases, and the diagnostic
level. There are two types of configuration files:
v The database manager configuration file for each DB2 UDB instance
v The database configuration file for each individual database.
The database manager configuration file is created when a DB2 UDB instance is
created. The parameters it contains affect system resources at the instance level,
independent of any one database that is part of that instance. Values for many of
these parameters can be changed from the system default values to improve
performance or increase capacity, depending on your system’s configuration.
There is one database manager configuration file for each client installation as well.
This file contains information about the client enabler for a specific workstation. A
subset of the parameters available for a server are applicable to the client.
Most of the parameters either affect the amount of system resources that will be
allocated to a single instance of the database manager, or they configure the setup
of the database manager and the different communications subsystems based on
environmental considerations. In addition, there are other parameters that serve
informative purposes only and cannot be changed. All of these parameters have
global applicability independent of any single database stored under that instance
of the database manager.
A database configuration file is created when a database is created, and resides where
that database resides. There is one configuration file per database. Its parameters
specify, among other things, the amount of resource to be allocated to that
database. Values for many of the parameters can be changed to improve
performance or increase capacity. Different changes may be required, depending on
the type of activity in a specific database.
Database Equivalent
object or concept physical object
Instance
Database manager
configuration parameters
Database
Database
configuration parameters
Related concepts:
v “Configuration parameter tuning” on page 316
Related tasks:
v “Configuring DB2 with configuration parameters” on page 317
Since the default values are oriented towards machines with relatively small
memory and dedicated as database servers, you may need to modify them if your
environment has:
v Large databases
v Large numbers of connections
v High performance requirements for a specific application
v Unique query or transaction loads or types
v Different machine configuration or usage.
Different types of applications and users have different response time requirements
and expectations. Applications could range from simple data entry screens to
strategic applications involving dozens of complex SQL statements accessing
dozens of tables per unit of work. For example, response time requirements could
vary considerably in a telephone customer service application versus a batch report
generation application.
Related concepts:
v “Configuration parameters” on page 315
Related tasks:
v “Configuring DB2 with configuration parameters” on page 317
Related reference:
v “Configuration parameters summary” on page 323
Attention: If you edit db2systm or SQLDBCON using a method other than those
provided by DB2, you may make the database unusable. We strongly recommend
that you do not change these files using methods other than those documented
and supported by DB2.
You may use one of the following methods to reset, update, and view
configuration parameters:
v Using the Control Center. The Configure Instance notebook can be used to set
the database manager configuration parameters on either a client or a server.
The Configure Database notebook can be used to alter the value of database
configuration parameters. The DB2 Control Center also provides the
Configuration Advisor to alter the value of configuration parameters. This
advisor generates values for parameters based on the responses you provide to a
set of questions, such as the workload and the type of transactions that run
against the database.
In a partitioned database environment, the SQLDBCON file exists for each database
partition. The Configure Database notebook will change the value on all
partitions if you launch the notebook from the database object in the tree view
of the Control Center. If you launch the notebook from a database partition
For some database manager configuration parameters, the database manager must
be stopped (db2stop) and then restarted (db2start) for the new parameter values
to take effect.
For some database parameters, changes will only take effect when the database is
reactivated. In these cases, all applications must first disconnect from the database.
(If the database was activated, then it must be deactivated and reactivated.) Then,
at the first new connect to the database, the changes will take effect.
Other parameters can be changed online; these are called configurable online
configuration parameters.
For clients, changes to the database manager configuration parameters take effect
the next time the client connects to a server.
Changing some database configuration parameters can influence the access plan
chosen by the SQL optimizer. After changing any of these parameters, you should
consider rebinding your applications to ensure the best access plan is being used
for your SQL statements. Any parameters that were modified online (for example,
by using the UPDATE DATABASE CONFIGURATION IMMEDIATE command)
will cause the SQL optimizer to choose new access plans for new SQL statements.
However, the SQL statement cache will not be purged of existing entries. To clear
the contents of the SQL cache, use the FLUSH PACKAGE CACHE statement.
While new parameter values may not be immediately effective, viewing the
parameter settings (using GET DATABASE MANAGER CONFIGURATION or GET
DATABASE CONFIGURATION commands) will always show the latest updates.
Viewing the parameter settings using the SHOW DETAIL clause on these
commands will show both the latest updates and the values in memory.
Related concepts:
v “Configuration parameters” on page 315
v “Configuration parameter tuning” on page 316
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
Given these workload characteristics, you could configure the system as shown in
the following table.
Daytime End of Day Daily Maintenance Weekly Maintenance
(05:00 - 20:00) (20:00 - 24:00) (24:00 - 05:00) (Sundays)
Database activity Heavy transaction Decision-support Load, backup, index Reorg and runstats
workload queries creation executed in buffer
pool 1
Buffer pool 1 (MB) 1000 500 500 2000
Buffer pool 2 (MB) 1000 500 500 200
Sort heap (MB) 0.1 20 200 200
Catalog cache (MB) 200 200 50 50
Package cache (MB) 800 200 200 200
Utility heap (MB) 0 0 1000 0
Diag level 1 3 4 4
You could then use the following scripts to transition the database from one
configuration to another. (These scripts can be scheduled to run at the appropriate
times.)
[Link]
[Link]
[Link]
Note that the order of operations differs from one script to the next. Operations
that reduce memory requirements should be completed before operations that
increase memory requirements; otherwise, the increases might fail.
All buffer pool resizings are followed by a commit operation, because buffer pool
changes are SQL operations and are part of the transaction.
The package cache is flushed after each reconfiguration to signal the optimizer that
the configuration has undergone a significant change, and that all existing dynamic
SQL access plans are now invalid.
For some database manager configuration parameters, the database manager must
be stopped (db2stop) and then restarted (db2start) for the new parameter values
to take effect. Other parameters can be changed online; these are called configurable
online configuration parameters. If you change the setting of a configurable online
database manager configuration parameter while you are attached to an instance,
the default behavior of the UPDATE DBM CFG command is to apply the change
immediately. If you do not want the change applied immediately, use the
DEFERRED option on the UPDATE DBM CFG command.
The column “Auto.” in the following table indicates whether the parameter
supports the AUTOMATIC keyword on the UPDATE DATABASE MANAGER
CONFIGURATION command. If you set a parameter to automatic, DB2 will
automatically adjust the parameter to reflect current resource requirements.
The columns “Token”, “Token Value”, and “Data Type” provide information that
you will need when calling the db2CfgGet or the db2CfgSet API. This information
includes configuration parameter identifiers, entries for the token element in the
db2CfgParam data structure, and data types for values that are passed to the
structure.
Table 41. Configurable Database Manager Configuration Parameters
Cfg. Perf. Token
Parameter Online Auto. Impact Token Value Data Type Additional Information
agent_stack_sz No No Low SQLF_KTN_AGENT_STACK_SZ 61 Uint16 “agent_stack_sz - Agent stack
size” on page 349
agentpri No No High SQLF_KTN_AGENTPRI 26 Sint16 “agentpri - Priority of agents” on
page 377
aslheapsz No No High SQLF_KTN_ASLHEAPSZ 15 Uint32 “aslheapsz - Application support
layer heap size” on page 358
audit_buf_sz No No High SQLF_KTN_AUDIT_BUF_SZ 312 Sint32 “audit_buf_sz - Audit buffer size”
on page 362
authentication1 No No Low SQLF_KTN_AUTHENTICATION 78 Uint16 “authentication - Authentication
type” on page 464
catalog_noauth Yes No None SQLF_KTN_CATALOG_NOAUTH 314 Uint16 “catalog_noauth - Cataloging
allowed without authority” on
page 465
| clnt_krb_plugin No No None SQLF_KTN_CLNT_KRB_PLUGIN 812 char(33) “clnt_krb_plugin - Client Kerberos
| plug-in” on page 466
clnt_pw_plugin No No None SQLF_KTN_CLNT_PW_PLUGIN 811 char(33) “clnt_pw_plugin - Client
userid-password plug-in” on page
466
comm_bandwidth Yes No Medium SQLF_KTN_COMM_BANDWIDTH 307 float “comm_bandwidth -
Communications bandwidth” on
page 456
conn_elapse Yes No Medium SQLF_KTN_CONN_ELAPSE 508 Uint16 “conn_elapse - Connection elapse
time” on page 443
cpuspeed Yes No Low2 SQLF_KTN_CPUSPEED 42 float “cpuspeed - CPU speed” on page
457
datalinks No No Low SQLF_KTN_DATALINKS 603 Sint16 “datalinks - Enable Data Links
support” on page 426
dft_account_str Yes No None SQLF_KTN_DFT_ACCOUNT_STR 28 char(25) “dft_account_str - Default
charge-back account” on page 458
dft_monswitches Yes No Medium SQLF_KTN_DFT_MONSWITCHES3 29 Uint16 “dft_monswitches - Default
v dft_mon_bufpool v SQLF_KTN_DFT_MON_BUFPOOL v 33 v Uint16 database system monitor switches”
on page 455
v dft_mon_lock v SQLF_KTN_DFT_MON_LOCK v 34 v Uint16
v dft_mon_sort v SQLF_KTN_DFT_MON_SORT v 35 v Uint16
v dft_mon_stmt v SQLF_KTN_DFT_MON_STMT v 31 v Uint16
v dft_mon_table v SQLF_KTN_DFT_MON_TABLE v 32 v Uint16
v dft_mon_timestamp v SQLF_KTN_DFT_MON_ v 36 v Uint16
v dft_mon_uow TIMESTAMP v 30 v Uint16
v SQLF_KTN_DFT_MON_UOW
dftdbpath Yes No None SQLF_KTN_DFTDBPATH 27 char(215) “dftdbpath - Default database
path” on page 467
diaglevel Yes No Low SQLF_KTN_DIAGLEVEL 64 Uint16 “diaglevel - Diagnostic error
capture level” on page 451
diagpath Yes No None SQLF_KTN_DIAGPATH 65 char(215) “diagpath - Diagnostic data
directory path” on page 452
dir_cache No No Medium SQLF_KTN_DIR_CACHE 40 Uint16 “dir_cache - Directory cache
support” on page 363
discover4 No No Medium SQLF_KTN_DISCOVER 304 Uint16 “discover - Discovery mode” on
page 441
discover_inst Yes No Low SQLF_KTN_DISCOVER_INST 308 Uint16 “discover_inst - Discover server
instance” on page 442
Notes:
1. Valid values (defined in sqlenv.h):
| SQL_AUTHENTICATION_SERVER (0)
| SQL_AUTHENTICATION_CLIENT (1)
| SQL_AUTHENTICATION_DCS (2)
| SQL_AUTHENTICATION_DCE (3)
| SQL_AUTHENTICATION_SVR_ENCRYPT (4)
| SQL_AUTHENTICATION_DCS_ENCRYPT (5)
| SQL_AUTHENTICATION_DCE_SVR_ENC (6)
| SQL_AUTHENTICATION_KERBEROS (7)
| SQL_AUTHENTICATION_KRB_SVR_ENC (8)
| SQL_AUTHENTICATION_GSSPLUGIN (9)
| SQL_AUTHENTICATION_GSS_SVR_ENC (10)
| SQL_AUTHENTICATION_DATAENC (11)
| SQL_AUTHENTICATION_DATAENC_CMP (12)
| SQL_AUTHENTICATION_NOT_SPEC (255)
2. The cpuspeed parameter can have a significant impact on performance, but you should use the default value, except in very specific circumstances,
as documented in the parameter description.
3. Bit 1 (xxxx xxx1): dft_mon_uow
Bit 2 (xxxx xx1x): dft_mon_stmt
Bit 3 (xxxx x1xx): dft_mon_table
Bit 4 (xxxx 1xxx): dft_mon_buffpool
Bit 5 (xxx1 xxxx): dft_mon_lock
Bit 6 (xx1x xxxx): dft_mon_sort
Bit 7 (x1xx xxxx): dft_mon_timestamp
4. Valid values (defined in sqlutil.h):
SQLF_DSCVR_KNOWN (1)
SQLF_DSCVR_SEARCH (2)
5. Valid values (defined in sqlutil.h):
SQLF_INX_REC_SYSTEM (0)
SQLF_INX_REC_REFERENCE (1)
6. Valid values (defined in sqlutil.h):
SQLF_TRUST_ALLCLNTS_NO (0)
SQLF_TRUST_ALLCLNTS_YES (1)
SQLF_TRUST_ALLCLNTS_DRDAONLY (2)
Notes:
1. Valid values (defined in sqlutil.h):
SQLF_NT_STANDALONE (0)
SQLF_NT_SERVER (1)
SQLF_NT_REQUESTOR (2)
SQLF_NT_STAND_REQ (3)
SQLF_NT_MPP (4)
SQLF_NT_SATELLITE (5)
For some database configuration parameters, changes will only take effect when
the database is reactivated. In these cases, all applications must first disconnect
from the database. (If the database was activated, then it must be deactivated and
reactivated.) The changes take effect at the next connection to the database. Other
parameters can be changed online; these are called configurable online configuration
parameters.
The column “Auto.” in the following table indicates whether the parameter
supports the AUTOMATIC keyword on the UPDATE DATABASE MANAGER
CONFIGURATION command. If you set a parameter to automatic, DB2 will
automatically adjust the parameter to reflect current resource requirements.
The columns “Token”, “Token Value”, and “Data Type” provide information that
you will need when calling the db2CfgGet or the db2CfgSet API. This information
Notes:
| 1. Default => Bit 1 on (xxxx xxxx xxxx xxx1): auto_maint
| Bit 2 off (xxxx xxxx xxxx xx0x): auto_db_backup
| Bit 3 on (xxxx xxxx xxxx x0xx): auto_tbl_maint
| Bit 4 on (xxxx xxxx xxxx 1xxx): auto_runstats
| Bit 5 off (xxxx xxxx xxx1 xxxx): auto_stats_prof
| Bit 6 off (xxxx xxxx xx0x xxxx): auto_prof_upd
| Bit 7 off (xxxx xxxx x0xx xxxx): auto_reorg
| 0 0 1 9
|
| Maximum => Bit 1 on (xxxx xxxx xxxx xxx1): auto_maint
| Bit 2 off (xxxx xxxx xxxx xx1x): auto_db_backup
| Bit 3 on (xxxx xxxx xxxx x1xx): auto_tbl_maint
| Bit 4 on (xxxx xxxx xxxx 1xxx): auto_runstats
| Bit 5 off (xxxx xxxx xxx1 xxxx): auto_stats_prof
| Bit 6 off (xxxx xxxx xx1x xxxx): auto_prof_upd
| Bit 7 off (xxxx xxxx x1xx xxxx): auto_reorg
| 0 0 7 F
2. Valid values (defined in sqlutil.h):
SQLF_INX_REC_SYSTEM (0)
SQLF_INX_REC_REFERENCE (1)
SQLF_INX_REC_RESTART (2)
3. Valid values (defined in sqlutil.h):
SQLF_LOGRETAIN_NO (0)
SQLF_LOGRETAIN_RECOVERY (1)
SQLF_LOGRETAIN_CAPTURE (2)
Notes:
1. char(17) on HP-UX and Solaris Operating Environment.
2. char(33) on HP-UX and Solaris Operating Environment.
Capacity management
There are a number of configuration parameters at both the database and database
manager levels that can impact the throughput on your system. These parameters
are categorized in the following groups:
v “Database shared memory”
v “Application shared memory” on page 346
v “Agent private memory” on page 349
v “Agent/application communication memory” on page 358
v “Database manager instance memory” on page 362
v “Locks” on page 367
v “I/O and storage” on page 370
v “Agents” on page 376
v “Stored procedures and user-defined functions” on page 386
This parameter is allocated out of the database shared memory, and is used to
cache system catalog information. In a partitioned database system, there is one
catalog cache for each database partition.
The use of the catalog cache can help improve the overall performance of:
v binding packages and compiling SQL statements
v operations that involve checking database-level privileges
v operations that involve checking execute privileges for routines
v applications that are connected to non-catalog nodes in a partitioned database
environment
Recommendation: Start with the default value and tune it by using the database
system monitor. When tuning this parameter, you should consider whether the
extra memory being reserved for the catalog cache might be more effective if it
was allocated for another purpose, such as the buffer pool or package cache.
Note: The catalog cache exists on all nodes in a partitioned database environment.
Since there is a local database configuration file for each node, each node’s
catalogcache_sz value defines the size of the local catalog cache. In order to
provide efficient caching and avoid overflow scenarios, you need to
explicitly set the catalogcache_sz value at each node and consider the
feasibility of possibly setting the catalogcache_sz on non-catalog nodes to be
smaller than that of the catalog node; keep in mind that information that is
required to be cached at non-catalog nodes will be retrieved from the
catalog node’s cache. Hence, a catalog cache at a non-catalog node is like a
subset of the information in the catalog cache at the catalog node.
In general, more cache space is required if a unit of work contains several dynamic
SQL statements or if you are binding packages that contain a large number of
static SQL statements.
This parameter specifies the amount of shared memory that is reserved for the
database shared memory region. If this amount is less than the amount calculated
from the individual parameters (for example, locklist, utility heap, bufferpools, and
so on), the larger amount will be used.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “Performance variables” on page 506
v “db2pd - Monitor and Troubleshoot DB2 Command” in the Command Reference
There is one database heap per database, and the database manager uses it on
behalf of all applications connected to the database. It contains control block
information for tables, indexes, table spaces, and buffer pools. It also contains
space for the log buffer (logbufsz) and temporary memory used by utilities.
Therefore, the size of the heap will be dependent on a large number of variables.
The control block information is kept in the heap until all applications disconnect
from the database.
| The minimum amount the database manager needs to get started is allocated at
| the first connection. The data area is expanded as needed until all the overflow
| memory area in the database shared memory is used.
You can use the database system monitor to track the highest amount of memory
that was used for the database heap, using the db_heap_top (maximum database
heap allocated) element.
Related reference:
v “logbufsz - Log buffer size” on page 342
v “db_heap_top - Maximum Database Heap Allocated monitor element” in the
System Monitor Guide and Reference
v “GET DATABASE CONFIGURATION Command” in the Command Reference
This parameter indicates the amount of storage that is allocated to the lock list.
There is one lock list per database and it contains the locks held by all applications
concurrently connected to the database. Locking is the mechanism that the
database manager uses to control concurrent access to data in the database by
multiple applications. Both rows and tables can be locked. The database manager
can also acquire locks for internal use.
This parameter can be changed online, but it can only be increased online, not
decreased. If you want to decrease the value of locklist, you will have to reactivate
the database.
On 32-bit platforms, each lock requires 36 or 72 bytes of the lock list, depending on
whether other locks are held on the object:
v 72 bytes are required to hold a lock on an object that has no other locks held on
it
v 36 bytes are required to record a lock on an object that has an existing lock held
on it.
On 64-bit platforms, each lock requires 56 or 112 bytes of the lock list, depending
on whether other locks are held on the object:
v 112 bytes are required to hold a lock on an object that has no other locks held on
it
v 56 bytes are required to record a lock on an object that has an existing lock held
on it.
When the percentage of the lock list used by one application reaches maxlocks, the
database manager will perform lock escalation, from row to table, for the locks
held by the application (described below). Although the escalation process itself
does not take much time, locking entire tables (versus individual rows) decreases
Once the lock list is full, performance can degrade since lock escalation will
generate more table locks and fewer row locks, thus reducing concurrency on
shared objects in the database. Additionally there might be more deadlocks
between applications (since they are all waiting on a limited number of table
locks), which will result in transactions being rolled back. Your application will
receive an SQLCODE of -912 when the maximum number of lock requests has
been reached for the database.
The following steps might help in determining the number of pages required for
your lock list:
1. Calculate a lower bound for the size of your lock list, using one of the following
calculations, depending on your environment:
a. (512 * x * maxappls) / 4096
b. with Concentrator enabled:
(512 * x * max_coordagents) / 4096
c. in a partitioned database with Concentrator enabled:
(512 * x * max_coordagents * number of database partitions) / 4096
where 512 is an estimate of the average number of locks per application and x
is the number of bytes required for each lock against an object that has an
existing lock (36 bytes on 32-bit platforms, 56 bytes on 64-bit platforms).
2. Calculate an upper bound for the size of your lock list:
(512 * y * maxappls) / 4096
where y is the number of bytes required for the first lock against an object (72
bytes on 32-bit platforms, 112 bytes on 64-bit platforms).
3. Estimate the amount of concurrency you will have against your data and based
on your expectations, choose an initial value for locklist that falls between the
upper and lower bounds that you have calculated.
4. Using the database system monitor, as described below, tune the value of this
parameter.
This information can help you validate or adjust the estimated number of locks per
application. In order to perform this validation, you will have to sample several
applications, noting that the monitor information is provided at a transaction level,
not an application level.
You should consider rebinding applications (using the REBIND command) after
changing this parameter.
Related reference:
v “maxlocks - Maximum percent of lock list before escalation” on page 369
v “maxappls - Maximum number of active applications” on page 381
v “lock_escals - Number of Lock Escalations monitor element” in the System
Monitor Guide and Reference
v “locks_held_top - Maximum Number of Locks Held monitor element” in the
System Monitor Guide and Reference
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “REBIND Command” in the Command Reference
This parameter allows you to specify the amount of the database heap (defined by
the dbheap parameter) to use as a buffer for log records before writing these records
to disk. The log records are written to disk when one of the following occurs:
v A transaction commits or a group of transactions commit, as defined by the
mincommit configuration parameter
v The log buffer is full
v As a result of some other internal database manager event.
This parameter must also be less than or equal to the dbheap parameter. Buffering
the log records will result in more efficient logging file I/O because the log records
will be written to disk less frequently and more log records will be written at each
time.
You can use the database system monitor to determine how much of the log buffer
space is used for a particular transaction (or unit of work). Refer to the
log_space_used (unit of work log space used) monitor element.
Related reference:
v “mincommit - Number of commits to group” on page 403
v “dbheap - Database heap” on page 339
v “uow_log_space_used - Unit of Work Log Space Used monitor element” in the
System Monitor Guide and Reference
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter is allocated out of the database shared memory, and is used for
caching of sections for static and dynamic SQL statements on a database. In a
partitioned database system, there is one package cache for each database partition.
Caching packages allows the database manager to reduce its internal overhead by
eliminating the need to access the system catalogs when reloading a package; or, in
the case of dynamic SQL, eliminating the need for compilation. Sections are kept in
the package cache until one of the following occurs:
v The database is shut down
v The package or dynamic SQL statement is invalidated
v The cache runs out of space.
This caching of the section for a static or dynamic SQL statement can improve
performance especially when the same statement is used multiple times by
applications connected to a database. This is particularly important in a transaction
processing application.
Recommendation: When tuning this parameter, you should consider whether the
extra memory being reserved for the package cache might be more effective if it
was allocated for another purpose, such as the buffer pool or catalog cache. For
this reason, you should use benchmarking techniques when tuning this parameter.
Tuning this parameter is particularly important when several sections are used
initially and then only a few are run repeatedly. If the cache is too large, memory
is wasted holding copies of the initial sections.
The following monitor elements can help you determine whether you should
adjust this configuration parameter:
v pkg_cache_lookups (package cache lookups)
v pkg_cache_inserts (package cache inserts)
v pkg_cache_size_top (package cache high water mark)
v pkg_cache_num_overflows (package cache overflows)
Note: The package cache is a working cache, so you cannot set this parameter to
zero. There must be sufficient memory allocated in this cache to hold all
sections of the SQL statements currently being executed. If there is more
space allocated than currently needed, then sections are cached. These
sections can simply be executed the next time they are needed without
having to load or compile them.
The limit specified by the pckcachesz parameter is a soft limit. This limit can
be exceeded, if required, if memory is still available in the database shared
set. You can use the pkg_cache_size_top monitor element to determine the
largest that the package cache has grown, and the pkg_cache_num_overflows
monitor element to determine how many times the limit specified by the
pckcachesz parameter has been exceeded.
Related reference:
v “maxappls - Maximum number of active applications” on page 381
v “pkg_cache_lookups - Package Cache Lookups monitor element” in the System
Monitor Guide and Reference
v “pkg_cache_inserts - Package Cache Inserts monitor element” in the System
Monitor Guide and Reference
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “pkg_cache_num_overflows - Package Cache Overflows monitor element” in the
System Monitor Guide and Reference
v “pkg_cache_size_top - Package Cache High Water Mark monitor element” in the
System Monitor Guide and Reference
This parameter represents a hard limit on the total amount of database shared
memory that can be used for sorting at any one time. When the total amount of
shared memory for active shared sorts reaches this limit, subsequent sorts will fail
(SQL0955C). If the value of sheapthres_shr is 0, the threshold for shared sort
memory will be equal to the value of the sheapthres database manager
configuration parameter, which is also used to represent the sort memory threshold
for private sorts. If the value of sheapthres_shr is non-zero, then this non-zero value
will be used for the shared sort memory threshold.
Related reference:
v “sheapthres - Sort heap threshold” on page 354
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter indicates the maximum amount of memory that can be used
simultaneously by the BACKUP, RESTORE, and LOAD (including load recovery)
utilities.
| Recommendation: Use the default value unless your utilities run out of space, in
| which case you should increase this value. If memory on your system is
| constrained, you might wish to lower the value of this parameter to limit the
| memory used by the database utilities. If the parameter is set too low and no more
| memory is available in the overflow area, you might not be able to concurrently
| run utilities. You should update this parameter dynamically as needed. For a small
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
The application control heap is required primarily for sharing information between
agents working on behalf of the same request. Usage of this heap is minimal for
non-partitioned databases when running queries with a degree of parallelism equal
to 1.
Recommendation: Initially, start with the default value. You might have to set the
value higher if you are running complex applications, if you have a system that
contains a large number of database partitions, or if you use declared temporary
tables. The amount of memory needed increases with the number of concurrently
active declared temporary tables. A declared temporary table with many columns
has a larger table descriptor size than a table with few columns, so having a large
number of columns in an application’s declared temporary tables also increases the
demand on the application control heap.
Related reference:
v “intra_parallel - Enable intra-partition parallelism” on page 449
v “applheapsz - Application heap size” on page 350
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “appgroup_mem_sz - Maximum size of application group memory set” on page
347
v “groupheap_ratio - Percent of memory for application group heap” on page 348
Recommendation: Retain the default value of this parameter unless you are
experiencing performance problems.
Related reference:
v “app_ctl_heap_sz - Application control heap size” on page 346
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “groupheap_ratio - Percent of memory for application group heap” on page 348
This parameter does not have any effect on a non-partitioned database with
concentrator OFF and intra-partition parallelism disabled.
Recommendation: Retain the default value of this parameter unless you are
experiencing performance problems.
Related reference:
v “app_ctl_heap_sz - Application control heap size” on page 346
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
The agent stack is the virtual memory that is allocated by DB2 for each agent. This
memory is committed when it is required to process an SQL statement. You can
use this parameter to optimize memory utilization of the server for a given set of
applications. More complex queries will use more stack space, compared to the
space used for simple queries.
This parameter is used to set the initial committed stack size for each agent in a
Windows environment. By default, each agent stack can grow up to the default
reserve stack size of 256 KB (64 4-KB pages). This limit is sufficient for most
database operations. However, when preparing a large SQL statement, the agent
can run out of stack space and the system will generate a stack overflow exception
(0xC000000D). When this happens, the server will shut down because the error is
non-recoverable.
The agent stack size can be increased by setting agent_stack_sz to a value larger
than the default reserve stack size of 64 pages. Note that the value for
agent_stack_sz, when larger than the default reserve stack size, is rounded by the
Windows operating system to the nearest multiple of 1 MB; setting the agent stack
You can change the default reserve stack size by using the db2hdr utility to change
the header information for the [Link] file. Changing the default reserve
stack size will affect all threads while changing agent_stack_sz only affects the stack
size for agents. The advantage of changing the default stack size using the db2hdr
utility is that it provides a better granularity, therefore allowing the stack size to be
set at the minimum required stack size. However, you will have to stop and restart
DB2 for a change to [Link] to take effect.
Recommendation: In most cases you should be able to use the default stack size.
Only if your environment includes many highly complex queries should you need
to increase the value of this parameter.
You might be able to reduce the stack size in order to make more address space
available to other clients, if your environment matches the following:
v Contains only simple applications (for example light OLTP), in which there are
never complex queries
v Requires a relatively large number of concurrent clients (for example, more than
100).
The agent stack size and the number of concurrent clients are inversely related: a
larger stack size reduces the potential number of concurrent clients that can be
running. This occurs because address space is limited on Windows platforms.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter defines the number of private memory pages available to be used
by the database manager on behalf of a specific agent or subagent.
Related reference:
v “app_ctl_heap_sz - Application control heap size” on page 346
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies the number of pages that the database server process will
reserve as private virtual memory, when a database manager instance is started
You should only change the value of this parameter if you want to commit more
memory to the database server. This action will save on allocation time. You
should be careful, however, that you do not set that value too high, as it can
impact the performance of non-DB2 applications.
Related reference:
v “priv_mem_thresh - Private memory threshold” on page 352
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter is used to determine the amount of unused agent private memory
that will be kept allocated, ready to be used by new agents that are started. It does
not apply to UNIX-based platforms.
A value of -1 will cause this parameter to use the value of the min_priv_mem
parameter.
Recommendation: When setting this parameter, you should consider the client
connection/disconnection patterns as well as the memory requirements of other
processes on the same machine.
If there is only a brief period during which many clients are concurrently
connected to the database, a high threshold will prevent unused memory from
being decommitted and made available to other processes. This case results in poor
memory management which can affect other processes which require memory.
If the number of concurrent clients is more uniform and there are frequent
fluctuations in this number, a high threshold will help to ensure memory is
available for the client processes and reduce the overhead to allocate and
deallocate memory.
This parameter specifies the maximum amount of memory that can be allocated
for the query heap. A query heap is used to store each query in the agent’s private
memory. The information for each query consists of the input and output SQLDA,
the statement text, the SQLCA, the package name, creator, section number, and
consistency token. This parameter is provided to ensure that an application does
not consume unnecessarily large amounts of virtual memory within an agent.
The query heap is also used for the memory allocated for blocking cursors. This
memory consists of a cursor control block and a fully resolved output SQLDA.
The initial query heap allocated will be the same size as the application support
layer heap, as specified by the aslheapsz parameter. The query heap size must be
greater than or equal to two (2), and must be greater than or equal to the aslheapsz
parameter. If this query heap is not large enough to handle a given request, it will
be reallocated to the size required by the request (not exceeding query_heap_sz). If
this new query heap is more than 1.5 times larger than aslheapsz, the query heap
will be reallocated to the size of aslheapsz when the query ends.
If you have very large LOBs, you might need to increase the value of this
parameter so the query heap will be large enough to accommodate those LOBs.
Related reference:
Private and shared sorts use memory from two different memory sources. The size
of the shared sort memory area is statically predetermined at the time of the first
connection to a database based on the value of sheapthres. The size of the private
sort memory area is unrestricted.
The sheapthres parameter is used differently for private and shared sorts:
v For private sorts, this parameter is an instance-wide soft limit on the total
amount of memory that can be consumed by private sorts at any given time.
When the total private-sort memory consumption for an instance reaches this
limit, the memory allocated for additional incoming private-sort requests will be
considerably reduced.
v For shared sorts, this parameter is a database-wide hard limit on the total
amount of memory consumed by shared sorts at any given time. When this limit
is reached, no further shared-sort memory requests will be allowed (until the
total shared-sort memory consumption falls below the limit specified by
sheapthres). (An alternate way to configure the shared-sort maximum value in
certain circumstances is to use the sheapthres_shr database configuration
parameter.)
Examples of operations that use the sort heap include: sorts, hash joins, dynamic
bitmaps (used for index ANDing and Star Joins), and operations where the table is
in memory.
Explicit definition of the threshold prevents the database manager from using
excessive amounts of memory for large numbers of sorts.
If you are doing private sorts and your system is not memory constrained, an ideal
value for this parameter can be calculated using the following steps:
1. Calculate the typical sort heap usage for each database:
(typical number of concurrent agents running against the database)
* (sortheap, as defined for that database)
2. Calculate the sum of the above results, which provides the total sort heap that
could be used under typical circumstances for all databases within the instance.
You should use benchmarking techniques to tune this parameter to find the proper
balance between sort performance and memory usage.
You can use the database system monitor to track the sort activity, using the post
threshold sorts (post_threshold_sorts) monitor element.
Related reference:
v “sortheap - Sort heap size” on page 355
v “post_threshold_sorts - Post Threshold Sorts monitor element” in the System
Monitor Guide and Reference
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “sheapthres_shr - Sort heap threshold for shared sorts” on page 344
This parameter defines the maximum number of private memory pages to be used
for private sorts, or the maximum number of shared memory pages to be used for
shared sorts. If the sort is a private sort, then this parameter affects agent private
memory. If the sort is a shared sort, then this parameter affects the database shared
memory. Each sort has a separate sort heap that is allocated as needed, by the
database manager. This sort heap is the area where data is sorted. If directed by
the optimizer, a smaller sort heap than the one specified by this parameter is
allocated using information provided by the optimizer.
Recommendation: When working with the sort heap, you should consider the
following:
v Appropriate indexes can minimize the use of the sort heap.
v Hash join buffers and dynamic bitmaps (used for index ANDing and Star Joins)
use sort heap memory. Increase the size of this parameter when these techniques
are used.
v Increase the size of this parameter when frequent large sorts are required.
v When increasing the value of this parameter, you should examine whether the
sheapthres parameter in the database manager configuration file also needs to be
adjusted.
v The sort heap size is used by the optimizer in determining access paths. You
should consider rebinding applications (using the REBIND command) after
changing this parameter.
Related reference:
v “sheapthres - Sort heap threshold” on page 354
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “REBIND Command” in the Command Reference
v “sheapthres_shr - Sort heap threshold for shared sorts” on page 344
This parameter indicates the maximum size of the heap used in collecting statistics
using the RUNSTATS command.
You should adjust this parameter based on the number of columns for which
statistics are being collected. Narrow tables, with relatively few columns, require
less memory for distribution statistics to be gathered. Wide tables, with many
columns, require significantly more memory. If you are gathering distribution
statistics for tables which are very wide and require a large statistics heap, you
might wish to collect the statistics during a period of low system activity so you
do not interfere with the memory requirements of other users.
Related reference:
v “num_freqvalues - Number of frequent values retained” on page 434
v “num_quantiles - Number of quantiles for columns” on page 435
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “RUNSTATS Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
The statement heap is used as a work space for the SQL compiler during
compilation of an SQL statement. This parameter specifies the size of this work
space.
This area does not stay permanently allocated, but is allocated and released for
every SQL statement handled. Note that for dynamic SQL statements, this work
area will be used during execution of your program; whereas, for static SQL
statements, it is used during the bind process but not during program execution.
Related reference:
v “sortheap - Sort heap size” on page 355
v “applheapsz - Application heap size” on page 350
The application support layer heap represents a communication buffer between the
local application and its associated agent. This buffer is allocated as shared
memory by each database manager agent that is started.
If the request to the database manager, or its associated reply, do not fit into the
buffer they will be split into two or more send-and-receive pairs. The size of this
buffer should be set to handle the majority of requests using a single
send-and-receive pair. The size of the request is based on the storage required to
hold:
v The input SQLDA
v All of the associated data in the SQLVARs
v The output SQLDA
v Other fields which do not generally exceed 250 bytes.
In addition to this communication buffer, this parameter is also used for two other
purposes:
v It is used to determine the I/O block size when a blocking cursor is opened.
This memory for blocked cursors is allocated out of the application’s private
address space, so you should determine the optimal amount of private memory
to allocate for each application program. If the database client cannot allocate
space for a blocking cursor out of an application’s private memory, a
non-blocking cursor will be opened.
The data sent from the local application is received by the database manager into a
set of contiguous memory allocated from the query heap. The aslheapsz parameter
is used to determine the initial size of the query heap (for both local and remote
clients). The maximum size of the query heap is defined by the query_heap_sz
parameter.
Use the following formula to calculate a minimum number of pages for aslheapsz:
aslheapsz >= ( sizeof(input SQLDA)
+ sizeof(each input SQLVAR)
+ sizeof(output SQLDA)
+ 250 ) / 4096
where sizeof(x) is the size of x in bytes that calculates the number of pages of a
given input or output value.
You should also consider the effect of this parameter on the number and potential
size of blocking cursors. Large row blocks might yield better performance if the
number or size of rows being transferred is large (for example, if the amount of
data is greater than 4 096 bytes). However, there is a trade-off in that larger record
blocks increase the size of the working set memory for each connection.
Larger record blocks might also cause more fetch requests than are actually
required by the application. You can control the number of fetch requests using the
OPTIMIZE FOR clause on the SELECT statement in your application.
Related reference:
v “query_heap_sz - Query heap size” on page 353
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
If this statement returns sqlcode SQL0419N, the database does not have
min_dec_div_3 support, or it is set to ″No″. If the statement returns 1.000,
min_dec_div_3 is set to ″Yes″.
2. min_dec_div_3 does not appear in the list of configuration keywords
when you run the following command: ? UPDATE DB CFG
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies the size of the communication buffer between remote
applications and their database agents on the database server. When a database
client requests a connection to a remote database, this communication buffer is
allocated on the client. On the database server, a communication buffer of 32 767
bytes is initially allocated, until a connection is established and the server can
determine the value of rqrioblk at the client. Once the server knows this value, it
will reallocate its communication buffer if the client’s buffer is not 32 767 bytes.
You should also consider the effect of this parameter on the number and potential
size of blocking cursors. Large row blocks might yield better performance if the
number or size of rows being transferred is large (for example, if the amount of
data is greater than 4 096 bytes). However, there is a trade-off in that larger record
blocks increase the size of the working set memory for each connection.
Larger record blocks might also cause more fetch requests than are actually
required by the application. You can control the number of fetch requests using the
OPTIMIZE FOR clause on the SELECT statement in your application.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
This parameter specifies the size of the buffer used when auditing the database.
The default value for this parameter is zero (0). If the value is zero (0), the audit
buffer is not used. If the value is greater than zero (0), space is allocated for the
audit buffer where the audit records will be placed when they are generated by the
audit facility. The value times 4 KB pages is the amount of space allocated for the
audit buffer. The audit buffer cannot be allocated dynamically; DB2 must be
stopped and then restarted before the new value for this parameter takes effect.
By changing this parameter from the default to some value larger than zero (0), the
audit facility writes records to disk asynchronously compared to the execution of
the statements generating the audit records. This improves DB2 performance over
leaving the parameter value at zero (0). The value of zero (0) means the audit
facility writes records to disk synchronously with (at the same time as) the
execution of the statements generating the audit records. The synchronous
operation during auditing decreases the performance of applications running in
DB2.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
By setting dir_cache to Yes the database, node and DCS directory files will be
cached in memory. The use of the directory cache reduces connect costs by
eliminating directory file I/O and minimizing the directory searches required to
retrieve directory information. There are two types of directory caches:
v An application directory cache that is allocated and used for each application
process on the machine at which the application is running.
v A server directory cache that is allocated and used for some of the internal
database manager processes.
For application directory caches, when an application issues its first connect, each
directory file is read and the information is cached in private memory for this
application. The cache is used by the application process on subsequent connect
requests and is maintained for the life of the application process. If a database is
not found in the application directory cache, the directory files are searched for the
information, but the cache is not updated. If the application modifies a directory
entry, the next connect within that application will cause the cache for this
application to be refreshed. The application directory cache for other applications
will not be refreshed. When the application process terminates, the cache is freed.
(To refresh the directory cache used by a command line processor session, issue a
db2 terminate command.)
Directory caching can also improve the performance of taking database system
monitor snapshots. In addition, you should explicitly reference the database name
on the snapshot call, instead of using database aliases.
Note: Errors might occur when performing snapshot calls if directory caching is
turned on and if databases are cataloged, uncataloged, created, or dropped
after the database manager is started.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the amount of memory that should be reserved for
instance management. This includes memory areas that describe the databases on
the instance.
| If you set this parameter to AUTOMATIC, DB2 will calculate the amount of
| instance memory needed for the current configuration. DB2 will also allocate some
| additional memory for an overflow buffer. The overflow buffer is used to satisfy
| peak memory requirements for any heap in the instance shared memory region
| whenever a heap exceeds its configured size. Other operations, such as dynamic
| configuration updates, also have access to this overflow buffer. The db2pd
Related reference:
v “maxagents - Maximum number of agents” on page 380
v “numdb - Maximum number of concurrently active databases including host
and iSeries databases” on page 460
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “db2pd - Monitor and Troubleshoot DB2 Command” in the Command Reference
This parameter determines the maximum size of the heap that is used by the Java
interpreter started to service Java DB2 stored procedures and UDFs.
There is one heap for each DB2 process (one for each agent or subagent on
UNIX-based platforms, and one for each instance on other platforms). There is one
heap for each fenced UDF and fenced stored procedure process. There is one heap
per agent (not including sub-agents) for trusted routines. There is one heap per
db2fmp process running a Java stored procedure. For multithreaded db2fmp
processes, multiple applications using threadsafe fenced routines are serviced from
a single heap. In all situations, only the agents or processes that run Java UDFs or
stored procedures ever allocate this memory. On partitioned database systems, the
same value is used at each partition.
Related reference:
v “jdk_path - Software Developer’s Kit for Java installation path” on page 459
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter determines the amount of the memory, in pages, to allocate for
database system monitor data. Memory is allocated from the monitor heap when
you perform database monitoring activities such as taking a snapshot, turning on a
monitor switch, resetting a monitor, or activating an event monitor.
A value of zero prevents the database manager from collecting database system
monitor data.
| If the available memory in this heap runs out and the overflow buffer has no
| unused memory, one of the following will occur:
v When the first application connects to the database for which this event monitor
is defined, an error message is written to the administration notification log.
v If an event monitor being started dynamically using the SET EVENT MONITOR
statement fails, an error code is returned to your application.
v If a monitor command or API subroutine fails, an error code is returned to your
application.
Related concepts:
v “Database system monitor memory requirements” in the System Monitor Guide
and Reference
Related reference:
Locks
The following parameters influence how locking is managed in your environment:
v “dlchktime - Time interval for checking deadlock”
v “locktimeout - Lock timeout” on page 368
v “maxlocks - Maximum percent of lock list before escalation” on page 369
See also “locklist - Maximum storage for lock list” on page 340.
A deadlock occurs when two or more applications connected to the same database
wait indefinitely for a resource. The waiting is never resolved because each
application is holding a resource that the other needs to continue.
The deadlock check interval defines the frequency at which the database manager
checks for deadlocks among all the applications connected to a database.
Notes:
1. In a partitioned database environment, this parameter applies to the catalog
node only.
2. In a partitioned database environment, a deadlock is not flagged until after the
second iteration.
Related reference:
v “locklist - Maximum storage for lock list” on page 340
This parameter specifies the number of seconds that an application will wait to
obtain a lock. This helps avoid global deadlocks for applications.
If you set this parameter to 0, locks are not waited for. In this situation, if no lock
is available at the time of the request, the application immediately receives a -911.
If you set this parameter to -1, lock timeout detection is turned off. In this
situation a lock will be waited for (if one is not available at the time of the request)
until either of the following:
v The lock is granted
v A deadlock occurs.
When working with Data Links Manager, if you see lock timeouts in the
administration notification log of the Data Links Manager (dlfm) instance, then you
should increase the value of locktimeout. You should also consider increasing the
value of locklist.
The value should be set to quickly detect waits that are occurring because of an
abnormal situation, such as a transaction that is stalled (possibly as a result of a
user leaving their workstation). You should set it high enough so valid lock
requests do not time-out because of peak workloads, during which time, there is
more waiting for locks.
You can use the database system monitor to help you track the number of times an
application (connection) experienced a lock timeout or that a database detected a
timeout situation for all applications that were connected.
High values of the lock_timeout (number of lock timeouts) monitor element can be
caused by:
v Too low a value for this configuration parameter.
v An application (transaction) that is holding locks for an extended period. You
can use the database system monitor to further investigate these applications.
v A concurrency problem, that could be caused by lock escalations (from row-level
to a table-level lock).
Related reference:
Lock escalation is the process of replacing row locks with table locks, reducing the
number of locks in the list. This parameter defines a percentage of the lock list
held by an application that must be filled before the database manager performs
escalation. When the number of locks held by any one application reaches this
percentage of the total lock list size, lock escalation will occur for the locks held by
that application. Lock escalation also occurs if the lock list runs out of space.
The database manager determines which locks to escalate by looking through the
lock list for the application and finding the table with the most row locks. If after
replacing these with a single table lock, the maxlocks value is no longer exceeded,
lock escalation will stop. If not, it will continue until the percentage of the lock list
held is below the value of maxlocks. The maxlocks parameter multiplied by the
maxappls parameter cannot be less than 100.
Where 2 is used to achieve twice the average and 100 represents the largest
percentage value allowed. If you have only a few applications that run
concurrently, you could use the following formula as an alternative to the first
formula:
maxlocks = 2 * 100 / (average number of applications running
concurrently)
One of the considerations when setting maxlocks is to use it in conjunction with the
size of the lock list (locklist). The actual limit of the number of locks held by an
application before lock escalation occurs is:
If maxlocks is set too low, lock escalation happens when there is still enough lock
space for other concurrent applications. If maxlocks is set too high, a few
applications can consume most of the lock space, and other applications will have
to perform lock escalation. The need for lock escalation in this case results in poor
concurrency.
You can use the database system monitor to help you track and tune this
configuration parameter.
Related reference:
v “locklist - Maximum storage for lock list” on page 340
v “maxappls - Maximum number of active applications” on page 381
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
Asynchronous page cleaners will write changed pages from the buffer pool (or the
buffer pools) to disk before the space in the buffer pool is required by a database
agent. As a result, database agents should not have to wait for changed pages to
be written out so that they might use the space in the buffer pool. This improves
overall performance of the database applications.
In a read-only (for example, query) environment, these page cleaners are not used.
Related reference:
v “num_iocleaners - Number of asynchronous page cleaners” on page 374
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
Recommendation: In many cases, you will want to explicitly specify the extent size
when you create the table space. Before choosing a value for this parameter, you
should understand how you would explicitly choose an extent size for the
CREATE TABLESPACE statement.
Related concepts:
v “Extent size” in the Administration Guide: Planning
Related reference:
v “dft_prefetch_sz - Default prefetch size” on page 372
v “CREATE TABLESPACE statement” in the SQL Reference, Volume 2
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| where the number of physical spindles defaults to 1 and can be specified through
| the DB2 registry variable DB2_PARALLEL_IO. This calculation is performed:
| v At database start-up time
| v When a table space is first created with AUTOMATIC prefetch size
| v When the number of containers for a table space changes through execution of
| an ALTER TABLESPACE statement
| v When the prefetch size for a table space is updated to be AUTOMATIC through
| execution of an ALTER TABLESPACE statement
| The AUTOMATIC state of the prefetch size can be turned on or off as soon as the
| prefetch size is updated manually through invocation of the ALTER TABLESPACE
| statement.
Recommendation: Using system monitoring tools, you can determine if your CPU
is idle while the system is waiting for I/O. Increasing the value of this parameter
can help if the table spaces being used do not have a prefetch size defined for
them.
This parameter provides the default for the entire database, and it might not be
suitable for all table spaces within the database. For example, a value of 32 might
be suitable for a table space with an extent size of 32 pages, but not suitable for a
table space with an extent size of 25 pages. Ideally, you should explicitly set the
prefetch size for each table space.
To help minimize I/O for table spaces defined with the default extent size
(dft_extent_sz), you should set this parameter as a factor or whole multiple of the
value of the dft_extent_sz parameter. For example, if the dft_extent_sz parameter is
32, you could set dft_prefetch_sz to 16 (a fraction of 32) or to 64 (a whole multiple of
32). If the prefetch size is a multiple of the extent size, the database manager might
perform I/O in parallel, if the following conditions are true:
v The extents being prefetched are on different physical devices
372 Administration Guide: Performance
v Multiple I/O servers are configured (num_ioservers).
Related reference:
v “ALTER TABLESPACE statement” in the SQL Reference, Volume 2
v “CREATE TABLESPACE statement” in the SQL Reference, Volume 2
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “System environment variables” on page 492
This parameter specifies the number of pages in each of the extended memory
segments in the database. This parameter is only used if your machine has more
real addressable memory than the maximum amount of virtual addressable
memory.
Related reference:
v “num_estore_segs - Number of extended storage memory segments” on page
373
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
Recommendation: Only use this parameter to establish the use of extended storage
memory segments if your platform environment has more memory than the
maximum address space and you wish to use this memory. When specifying the
number of segments, you should also consider the size of the each of the segments
by reviewing and modifying the estore_seg_sz parameter.
Related reference:
v “estore_seg_sz - Extended storage memory segment size” on page 373
v “ALTER BUFFERPOOL statement” in the SQL Reference, Volume 2
v “CREATE BUFFERPOOL statement” in the SQL Reference, Volume 2
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter allows you to specify the number of asynchronous page cleaners
for a database. These page cleaners write changed pages from the buffer pool to
disk before the space in the buffer pool is required by a database agent. As a
result, database agents should not have to wait for changed pages to be written
out so that they might use the space in the buffer pool. This improves overall
performance of the database applications.
If you set the parameter to zero (0), no page cleaners are started and as a result,
the database agents will perform all of the page writes from the buffer pool to
disk. This parameter can have a significant performance impact on a database
stored across many physical storage devices, since in this case there is a greater
chance that one of the devices will be idle. If no page cleaners are configured, your
applications might encounter periodic log full conditions.
If the applications for a database primarily consist of transactions that update data,
an increase in the number of cleaners will speed up performance. Increasing the
page cleaners will also decrease recovery time from soft failures, such as power
outages, because the contents of the database on disk will be more up-to-date at
any given time.
Recommendation: Consider the following factors when setting the value for this
parameter:
v Application type
– If it is a query-only database that will not have updates, set this parameter to
be zero (0). The exception would be if the query work load results in many
TEMP tables being created (you can determine this by using the explain
utility).
– If transactions are run against the database, set this parameter to be between
one and the number of physical storage devices used for the database.
v Workload
Environments with high update transaction rates might require more page
cleaners to be configured.
v Buffer pool sizes
You can use the database system monitor to help you tune this configuration
parameter using information from the event monitor about write activity from a
buffer pool:
v The parameter can be reduced if both of the following conditions are true:
– pool_data_writes is approximately equal to pool_async_data_writes
– pool_index_writes is approximately equal to pool_async_index_writes.
v The parameter should be increased if either of the following conditions are true:
– pool_data_writes is much greater than pool_async_data_writes
– pool_index_writes is much greater than pool_async_index_writes.
Related reference:
v “chngpgs_thresh - Changed pages threshold” on page 370
v “pool_data_writes - Buffer Pool Data Writes monitor element” in the System
Monitor Guide and Reference
v “pool_index_writes - Buffer Pool Index Writes monitor element” in the System
Monitor Guide and Reference
v “pool_async_data_writes - Buffer Pool Asynchronous Data Writes monitor
element” in the System Monitor Guide and Reference
v “pool_async_index_writes - Buffer Pool Asynchronous Index Writes monitor
element” in the System Monitor Guide and Reference
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
I/O servers are used on behalf of the database agents to perform prefetch I/O and
asynchronous I/O by utilities such as backup and restore. This parameter specifies
the number of I/O servers for a database. No more than this number of I/Os for
prefetching and utilities can be in progress for a database at any time. An I/O
server waits while an I/O operation that it initiated is in progress. Non-prefetch
I/Os are scheduled directly from the database agents and as a result are not
constrained by num_ioservers.
Recommendation: In order to fully exploit all the I/O devices in the system, a
good value to use is generally one or two more than the number of physical
devices on which the database resides. It is better to configure additional I/O
servers, since there is minimal overhead associated with each I/O server and any
unused I/O servers will remain idle.
Related reference:
This parameter, which only applies to SMS table spaces, indicates the number of
containers that will be created within the default table spaces. This parameter will
show the information used when you created your database, whether it was
specified explicitly or implicitly on the CREATE DATABASE command. The
CREATE TABLESPACE statement does not use this parameter in any way.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
The database manager can monitor I/O and if sequential page reading is occurring
the database manager can activate I/O prefetching. This type of sequential prefetch
is known as sequential detection. You can use the seqdetect configuration parameter
to control whether the database manager should perform sequential detection.
If this parameter is set to No, prefetching takes place only if the database manager
knows it will be useful, for example table sorts, table scans, or list prefetch.
Recommendation: In most cases, you should use the default value for this
parameter. Try turning sequential detection off, only if other tuning efforts were
unable to correct serious query performance problems.
Related reference:
v “dft_prefetch_sz - Default prefetch size” on page 372
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
Agents
The following parameters can influence the number of applications that can be run
concurrently and achieve optimal performance:
v “agentpri - Priority of agents” on page 377
v “avg_appls - Average number of active applications” on page 378
This parameter controls the priority given both to all agents, and to other database
manager instance processes and threads, by the operating system scheduler. In a
partitioned database environment, this also includes both coordinating and
subagents, the parallel system controllers, and the FCM daemons. This priority
determines how CPU time is given to the DB2 processes, agents, and threads
relative to the other processes and threads running on the machine. When the
parameter is set to -1, no special action is taken and the database manager is
scheduled in the normal way that the operating system schedules all processes and
threads. When the parameter is set to a value other than -1, the database manager
will create its processes and threads with a static priority set to the value of the
parameter. Therefore, this parameter allows you to control the priority with which
the database manager processes and threads will execute on your machine.
You can use this parameter to increase database manager throughput. The values
for setting this parameter are dependent on the operating system on which the
database manager is running. For example, in a UNIX-based environment,
numerically low values yield high priorities. When the parameter is set to a value
between 41 and 125, the database manager creates its agents with a UNIX static
priority set to the value of the parameter. This is important in UNIX-based
environments because numerically low values yield high priorities for the database
manager, but other processes (including applications and users) might experience
delays because they cannot obtain enough CPU time. You should balance the
setting of this parameter with the other activity expected on the machine.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter is used by the SQL optimizer to help estimate how much buffer
pool will be available at run-time for the access plan chosen.
When setting this parameter, you should estimate the number of complex query
applications that typically use the database. This estimate should exclude all light
OLTP applications. If you have trouble estimating this number, you can multiply
the following:
v An average number of all applications running against your database. The
database system monitor can provide information about the number of
applications at any given time and using a sampling technique, you can
calculate an average over a period of time. The information from the database
system monitor includes both OLTP and non-OLTP applications.
v Your estimate of the percentage of complex query applications.
As with adjusting other configuration parameters that affect the optimizer, you
should adjust this parameter in small increments. This allows you to minimize
path selection differences.
Related reference:
v “maxappls - Maximum number of active applications” on page 381
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
When the Concentrator is off, this parameter indicates the maximum number of
client connections allowed per partition. The Concentrator is off when
max_connections is equal to max_coordagents. The Concentrator is on when
max_connections is greater than max_coordagents.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
One coordinating agent is acquired for each local or remote application that
connects to a database or attaches to an instance. Requests that require an instance
attachment include CREATE DATABASE, DROP DATABASE, and Database System
Monitor commands.
When the Concentrator is on, that is, when max_connections is greater than
max_coordagents, there might be more connections than coordinator agents to
service them. An application is in an active state only if there is a coordinator
agent servicing it. Otherwise, the application is in an inactive state. Requests from
an active application will be serviced by the database coordinator agent (and
subagents in SMP or MPP configurations). Requests from an inactive application
will be queued until a database coordinator agent is assigned to service the
application, when the application becomes active. As a result, this parameter can
be used to control the load on the system.
Related reference:
v “num_initagents - Initial number of agents in pool” on page 385
v “num_poolagents - Agent pool size” on page 385
v “intra_parallel - Enable intra-partition parallelism” on page 449
v “maxagents - Maximum number of agents” on page 380
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
Recommendation: The value of maxagents should be at least the sum of the values
for maxappls in each database allowed to be accessed concurrently. If the number of
databases is greater than the numdb parameter, then the safest course is to use the
product of numdb with the largest value for maxappls.
Each additional agent requires some resource overhead that is allocated at the time
the database manager is started.
Related reference:
v “max_coordagents - Maximum number of coordinating agents” on page 379
v “num_poolagents - Agent pool size” on page 385
v “maxcagents - Maximum number of concurrent agents” on page 383
v “fenced_pool - Maximum number of fenced processes” on page 386
v “maxappls - Maximum number of active applications” on page 381
v “min_priv_mem - Minimum committed private memory” on page 351
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the maximum number of concurrent applications that can
be connected (both local and remote) to a database. Since each application that
Setting maxappls to automatic has the effect of allowing any number of connected
applications. DB2 will dynamically allocate the resources it needs to support new
applications.
If you do not want to set this parameter to automatic, the value of this parameter
must be equal to or greater than the sum of the connected applications, plus the
number of these same applications that might be concurrently in the process of
completing a two-phase commit or rollback. Then add to this sum the anticipated
number of indoubt transactions that might exist at any one time.
As more applications use the Data Links Manager, the value of maxappls should be
increased. Use the following formula to compute the value you need:
<maxappls> = 5 * (number of nodes) + (peak number of active applications
using Data Links Manager)
Related tasks:
v “Manually resolving indoubt transactions” in the Administration Guide: Planning
Related reference:
v “max_coordagents - Maximum number of coordinating agents” on page 379
v “maxagents - Maximum number of agents” on page 380
v “locklist - Maximum storage for lock list” on page 340
v “maxlocks - Maximum percent of lock list before escalation” on page 369
v “avg_appls - Average number of active applications” on page 378
This parameter does not limit the number of applications that can have
connections to a database. It only limits the number of database manager agents
that can be processed concurrently by the database manager at any one time,
thereby limiting the usage of system resources during times of peak processing.
Recommendation: In most cases the default value for this parameter will be
acceptable. In cases where the high concurrency of applications is causing
problems, you can use benchmark testing to tune this parameter to optimize the
performance of the database.
Related reference:
v “max_coordagents - Maximum number of coordinating agents” on page 379
v “maxagents - Maximum number of agents” on page 380
v “maxappls - Maximum number of active applications” on page 381
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the maximum number of file handles that can be open for
each database agent. If opening a file causes this value to be exceeded, some files
in use by this agent are closed. If maxfilop is too small, the overhead of opening
and closing files so as not to exceed this limit will become excessive and might
degrade performance.
Both SMS table spaces and DMS table space file containers are treated as files in
the database manager’s interaction with the operating system, and file handles are
required. More files are generally used by SMS table spaces compared to the
number of containers used for a DMS file table space. Therefore, if you are using
SMS table spaces, you will need a larger value for this parameter compared to
what you would require for DMS file table spaces.
You can also use this parameter to ensure that the overall total of file handles used
by the database manager does not exceed the operating system limit by limiting
the number of handles per agent to a specific number; the actual number will vary
depending on the number of agents running concurrently.
Related reference:
v “maxappls - Maximum number of active applications” on page 381
v “maxtotfilop - Maximum total files open” on page 384
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter defines the maximum number of files that can be opened by all
agents and other threads executing in a single database manager instance. If
opening a file causes this value to be exceeded, an error is returned to your
application.
If a new database is created, you should re-evaluate the value for this parameter.
Related reference:
v “maxfilop - Maximum database files open per application” on page 383
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter determines the initial number of idle agents that are created in the
agent pool at DB2START time.
Related reference:
v “max_coordagents - Maximum number of coordinating agents” on page 379
v “num_poolagents - Agent pool size” on page 385
v “maxagents - Maximum number of agents” on page 380
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
When the Concentrator is on, that is, when max_connections is greater than
max_coordagents, agents will always be returned to the pool, regardless of the value
of this parameter. Based on the system load and the time agents remain idle in the
pool, agents might terminate themselves, as necessary, to reduce the size of the idle
pool to the configured parameter value.
Except when the Concentrator is on, if the value of this parameter is 0, agents will
be created as needed, and will terminate once they finish executing their current
request.
Related reference:
v “num_initagents - Initial number of agents in pool” on page 385
v “max_coordagents - Maximum number of coordinating agents” on page 379
v “max_querydegree - Maximum query degree of parallelism” on page 450
v “maxagents - Maximum number of agents” on page 380
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
If the parameter is set to −1, the maximum number of cached db2fmp processes
will be the same as the value set in the max_coordagents parameter.
If you find that the default value is not appropriate for your environment because
an inappropriate amount of system resource is being given to db2fmp processes
and is affecting performance of the database manager, the following might be
useful in providing a starting point for tuning this parameter:
fenced_pool = # of applications allowed to make stored procedure and
UDF calls at one time
If keepfenced is set to yes, then each db2fmp process that is created in the cache pool
will continue to exist and use system resources even after the fenced routine call
has been processed and returned to the agent.
If keepfenced is set to no, then nonthreaded db2fmp processes will terminate when
they complete execution, and there is no cache pool. Multithreaded db2fmp
processes will continue to exist, but no threads will be pooled in these processes.
This means that even when keepfenced is set no you can have one threaded C
db2fmp process and one threaded Java db2fmp process on your system.
Related reference:
v “max_coordagents - Maximum number of coordinating agents” on page 379
v “maxagents - Maximum number of agents” on page 380
v “keepfenced - Keep fenced process” on page 388
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter indicates whether or not a fenced mode process is kept after a
fenced mode routine call is complete. Fenced mode processes are created as
separate system entities in order to isolate user-written fenced mode code from the
database manager agent process. This parameter is only applicable on database
servers.
If keepfenced is set to no, and the routine being executed is not threadsafe, a new
fenced mode process is created and destroyed for each fenced mode invocation. If
keepfenced is set to no, and the routine being executed is threadsafe, the fenced
mode process persists, but the thread created for the call is terminated. If keepfenced
is set to yes, a fenced mode process or thread is reused for subsequent fenced
mode calls. When the database manager is stopped, all outstanding fenced mode
processes and threads will be terminated.
Setting this parameter to yes will result in additional system resources being
consumed by the database manager for each fenced mode process that is activated,
up to the value contained in the fenced_pool parameter. A new process is only
created when no existing fenced mode process is available to process a subsequent
fenced routine invocation. This parameter is ignored if fenced_pool is set to 0.
This parameter indicates the initial number of nonthreaded, idle db2fmp processes
that are created in the db2fmp pool at DB2START time. Setting this parameter will
reduce the initial startup time for running non-threadsafe C and Cobol routines.
This parameter is ignored if keepfenced is not specified.
It is much more important to set fenced_pool to an appropriate size for your system
than to start up a number of db2fmp processes at DB2START time.
Related reference:
v “keepfenced - Keep fenced process” on page 388
v “fenced_pool - Maximum number of fenced processes” on page 386
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter defines the size of each primary and secondary log file. The size of
these log files limits the number of log records that can be written to them before
they become full and a new log file is required.
The use of primary and secondary log files as well as the action taken when a log
file becomes full are dependent on the type of logging that is being performed:
v Circular logging
A primary log file can be reused when the changes recorded in it have been
committed. If the log file size is small and applications have processed a large
number of changes to the database without committing the changes, a primary
log file can quickly become full. If all primary log files become full, the database
manager will allocate secondary log files to hold the new log records.
v Log retention logging
When a primary log file is full, the log is archived and a new primary log file is
allocated.
Recommendation: You must balance the size of the log files with the number of
primary log files:
v The value of the logfilsiz should be increased if the database has a large number
of update, delete, or insert transactions running against it which will cause the
log file to become full very quickly.
Note: The upper limit of log file size, combined with the upper limit of the
number of log files (logprimary + logsecond), gives an upper limit of 256 GB
of active log space.
If you are using log retention, the current active log file is closed and truncated
when the last application disconnects from a database. When the next connection
to the database occurs, the next log file is used. Therefore, if you understand the
logging requirements of your concurrent applications, you might be able to
determine a log file size that will not allocate excessive amounts of wasted space.
Related reference:
v “logprimary - Number of primary log files” on page 391
v “logsecond - Number of secondary log files” on page 393
v “softmax - Recovery range and soft checkpoint interval” on page 405
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter contains the name of the log file that is currently active.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
This parameter contains the current path being used for logging purposes. You
cannot change this parameter directly as it is set by the database manager after a
change to the newlogpath parameter becomes effective.
When a database is created, the recovery log file for it is created in a subdirectory
of the directory containing the database. The default is a subdirectory named
SQLOGDIR under the directory created for the database.
Related reference:
v “newlogpath - Change the database log path” on page 396
v “GET DATABASE CONFIGURATION Command” in the Command Reference
The primary log files establish a fixed amount of storage allocated to the recovery
log files. This parameter allows you to specify the number of primary log files to
be preallocated.
Under circular logging, the primary logs are used repeatedly in sequence. That is,
when a log is full, the next primary log in the sequence is used if it is available. A
log is considered available if all units of work with log records in it have been
committed or rolled-back. If the next primary log in sequence is not available, then
a secondary log is allocated and used. Additional secondary logs are allocated and
used until the next primary log in the sequence becomes available or the limit
imposed by the logsecond parameter is reached. These secondary log files are
dynamically deallocated as they are no longer needed by the database manager.
The number of primary and secondary log files must comply with the following:
v If logsecond has a value of -1, logprimary <= 256.
v If logsecond does not have a value of -1, (logprimary + logsecond) <= 256.
Increasing this value will increase the disk requirements for the logs because the
primary log files are preallocated during the very first connection to the database.
If you find that secondary log files are frequently being allocated, you might be
able to improve system performance by increasing the log file size (logfilsiz) or by
increasing the number of primary log files.
For databases that are not frequently accessed, in order to save disk storage, set the
parameter to 2. For databases enabled for roll-forward recovery, set the parameter
larger to avoid the overhead of allocating new logs almost immediately.
Related reference:
v “logfilsiz - Size of log files” on page 390
v “logsecond - Number of secondary log files” on page 393
v “logretain - Log retain enable” on page 403
v “userexit - User exit enable” on page 406
v “sec_log_used_top - Maximum Secondary Log Space Used monitor element” in
the System Monitor Guide and Reference
v “tot_log_used_top - Maximum Total Log Space Used monitor element” in the
System Monitor Guide and Reference
v “sec_logs_allocated - Secondary Logs Allocated Currently monitor element” in
the System Monitor Guide and Reference
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies the number of secondary log files that are created and
used for recovery log files (only as needed). When the primary log files become
full, the secondary log files (of size logfilsiz) are allocated one at a time as needed,
up to a maximum number as controlled by this parameter. An error code will be
returned to the application, and the database will be shut down, if more secondary
log files are required than are allowed by this parameter.
If you set logsecond to -1, the database is configured with infinite active log space.
There is no limit on the size or the number of in-flight transactions running on the
database. If you set logsecond to -1, you still use the logprimary and logfilsiz
configuration parameters to specify how many log files DB2 should keep in the
active log path. If DB2 needs to read log data from a log file, but the file is not in
the active log path, DB2 will invoke the userexit program to retrieve the log file
from the archive to the active log path. (DB2 will retrieve the files to the overflow
log path, if you have configured one.) Once the log file is retrieved, DB2 will cache
If your log path is a raw device, you must configure the overflowlogpath
configuration parameter in order to set logsecond to -1.
By setting logsecond to -1, you will have no limit on the size of the unit of work or
the number of concurrent units of work. However, rollback (both at the savepoint
level and at the unit of work level) could be very slow due to the need to retrieve
log files from the archive. Crash recovery could also be very slow for the same
reason. DB2 will write a message to the administration notification log to warn you
that the current set of active units of work has exceeded the primary log files. This
is an indication that rollback or crash recovery could be extremely slow.
Recommendation: Use secondary log files for databases that have periodic needs
for large amounts of log space. For example, an application that is run once a
month might require log space beyond that provided by the primary log files.
Since secondary log files do not require permanent file space they are
advantageous in this type of situation.
Related reference:
v “logfilsiz - Size of log files” on page 390
v “logprimary - Number of primary log files” on page 391
v “logretain - Log retain enable” on page 403
v “userexit - User exit enable” on page 406
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “overflowlogpath - Overflow log path” on page 398
If the value is not 0, this parameter indicates the percentage of active log space that
can be consumed by one transaction.
If the value is set to 0, there is no limit regarding how much space (as a percentage
of total active log space) one single transaction can consume. This was the
behavior of transactions prior to Version 8.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter allows you to specify a string of up to 242 bytes for the mirror log
path. The string must point to a path name, and it must be a fully qualified path
name, not a relative path name.
If mirrorlogpath is configured, DB2 will create active log files in both the log path
and the mirror log path. All log data will be written to both paths. The mirror log
path has a duplicated set of active log files, such that if there is a disk error or
human error that destroys active log files on one of the paths, the database can still
function.
If the mirror log path is changed, there might be log files in the old mirror log
path. These log files might not have been archived, so you might need to archive
these log files manually. Also, if you are running replication on this database,
replication might still need the log files from before the log path change. If the
database is configured with the User Exit Enable (userexit) database configuration
parameter set to Yes, and if all the log files have been archived either by DB2
automatically or by yourself manually, then DB2 will be able to retrieve the log
files to complete the replication process. Otherwise, you can copy the files from the
old mirror log path to the new mirror log path.
If logpath or newlogpath specifies a raw device as the location where the log files are
stored, mirror logging, as indicated by mirrorlogpath, is not allowed. If logpath or
newlogpath specifies a file path as the location where the log files are stored, mirror
logging is allowed and mirrorlogpath must also specify a file path.
Recommendation: Just like the log files, the mirror log files should be on a
physical disk that does not have high I/O.
It is strongly recommended that this path be on a separate device than the primary
log path.
You can use the database system monitor to track the number of I/Os related to
database logging.
The following data elements return the amount of I/O activity related to database
logging. You can use an operating system monitor tool to collect information about
other disk I/O activity, then compare the two types of I/O activity.
v log_reads (number of log pages read)
v log_writes (number of log pages written).
Related reference:
v “logpath - Location of log files” on page 391
v “newlogpath - Change the database log path” on page 396
This parameter allows you to specify a string of up to 242 bytes to change the
location where the log files are stored. The string can point to either a path name
or to a raw device. If the string points to a path name, it must be a fully qualified
path name, not a relative path name.
If you want to use replication, and your log path is a raw device, the
overflowlogpath configuration parameter must be configured.
Note: You must have Windows NT Version 4.0 with Service Pack 3 or later
installed to be able to write logs to a device.
v On UNIX-based platforms, /dev/rdblog8
Note: You can only specify a device on AIX, Windows 2000, Windows NT, Solaris
Operating Environment, HP-UX, and Linux platforms.
The new setting does not become the value of logpath until both of the following
occur:
v The database is in a consistent state, as indicated by the database_consistent
parameter.
v All users are disconnected from the database
When the first new connection is made to the database, the database manager will
move the logs to the new location specified by logpath.
There might be log files in the old log path. These log files might not have been
archived. You might need to archive these log files manually. Also, if you are
running replication on this database, replication might still need the log files from
before the log path change. If the database is configured with the User Exit Enable
(userexit) database configuration parameter set to Yes, and if all the log files have
been archived either by DB2 automatically or by yourself manually, then DB2 will
If logpath or newlogpath specifies a raw device as the location where the log files are
stored, mirror logging, as indicated by mirrorlogpath, is not allowed. If logpath or
newlogpath specifies a file path as the location where the log files are stored, mirror
logging is allowed and mirrorlogpath must also specify a file path.
Recommendation: Ideally, the log files will be on a physical disk which does not
have high I/O. For instance, avoid putting the logs on the same disk as the
operating system or high volume databases. This will allow for efficient logging
activity with a minimum of overhead such as waiting for I/O.
You can use the database system monitor to track the number of I/Os related to
database logging.
The monitor elements log_reads (number of log pages read) and log_writes (number
of log pages written) return the amount of I/O activity related to database logging.
You can use an operating system monitor tool to collect information about other
disk I/O activity, then compare the two types of I/O activity.
Related reference:
v “logpath - Location of log files” on page 391
v “database_consistent - Database is consistent” on page 429
v “log_reads - Number of Log Pages Read monitor element” in the System Monitor
Guide and Reference
v “log_writes - Number of Log Pages Written monitor element” in the System
Monitor Guide and Reference
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
If the value is not 0, this parameter indicates the number of active log files that one
active transaction is allowed to span.
If the value is set to 0, there is no limit to how many log files one single
transaction can span. This was the behavior of transactions prior to Version 8.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “max_log - Maximum log per transaction” on page 394
This parameter can be used for several functions, depending on your logging
requirements.
v This parameter allows you to specify a location for DB2 to find log files that are
needed for a rollforward operation. It is similar to the OVERFLOW LOG PATH
option on the ROLLFORWARD command. Instead of always specifying
OVERFLOW LOG PATH on every ROLLFORWARD command, you can set this
configuration parameter once. However, if both are used, the OVERFLOW LOG
PATH option will overwrite the overflowlogpath configuration parameter, for that
particular rollforward operation.
v If logsecond is set to -1, overflowlogpath allows you to specify a directory for DB2
to store active log files retrieved from the archive. (Active log files have to be
retrieved for rollback operations if they are no longer in the active log path).
Without overflowlogpath, DB2 will retrieve the log files into the active log path.
Using overflowlogpath allows you to provide additional resource for DB2 to store
the retrieved log files. The benefit includes spreading the I/O cost to different
disks, and allowing more log files to be stored in the active log path.
v If you need to use the db2ReadLog API (prior to DB2 V8, db2ReadLog was
called sqlurlog) for replication, for example, overflowlogpath allows you to specify
a location for DB2 to search for log files that are needed for this API. If the log
file is not found (in either the active log path or the overflow log path) and the
database is configured with userexit enabled, DB2 will retrieve the log file.
overflowlogpath also allows you to specify a directory for DB2 to store the log
files retrieved. The benefit comes from reducing the I/O cost on the active log
path and allowing more log files to be stored in the active log path.
v If you have configured a raw device for the active log path, overflowlogpath must
be configured if you want to set logsecond to -1, or if you want to use the
db2ReadLog API.
To set overflowlogpath, specify a string of up to 242 bytes. The string must point to a
path name, and it must be a fully qualified path name, not a relative path name.
The path name must be a directory, not a raw device.
Related reference:
v “logsecond - Number of secondary log files” on page 393
v “db2ReadLog - Asynchronous Read Log” in the Administrative API Reference
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “ROLLFORWARD DATABASE Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the number of seconds to wait after a failed archive
| attempt before trying to archive the log file again. Subsequent retries will only take
| affect if the value of the numarchretry database configuration parameter is at least 1.
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This configuration parameter can be set to prevent disk full errors from being
generated when DB2 cannot create a new log file in the active log path. Instead,
DB2 will attempt to create the log file every five minutes until it succeeds. After
If blk_log_dsk_ful is set to no, then a transaction that receives a log disk full error
will fail and will be rolled back. In some situations, the database will come down if
a transaction causes a log disk full error.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies a path to which DB2 will try to archive log files if the log
| files cannot be archived to either the primary or the secondary (if set) archive
| destinations because of a media problem affecting those destinations. This specified
| path must reference a disk.
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the media type of the primary destination for archived
| logs.
| Related concepts:
| v Appendix E, “Cross-node recovery with the db2adutl command and the
| logarchopt1 and vendoropt database configuration parameters,” on page 589
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the media type of the secondary destination for archived
| logs. If this path is specified, log files will be archived to both this destination and
| the destination specified by the logarchmeth1 database configuration parameter.
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the options field for the primary destination for archived
| logs (if required).
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the options field for the secondary destination for
| archived logs (if required).
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
If logretain is set to Recovery or userexit is set to Yes, the active log files will be
retained and become online archive log files for use in roll-forward recovery. This
is called log retention logging.
After logretain is set to Recovery or userexit is set to Yes (or both), you must make a
full backup of the database. This state is indicated by the backup_pending flag
parameter.
Related reference:
v “log_retain_status - Log retain status indicator” on page 429
v “userexit - User exit enable” on page 406
v “backup_pending - Backup pending indicator” on page 429
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter allows you to delay the writing of log records to disk until a
minimum number of commits have been performed. This delay can help reduce
the database manager overhead associated with writing log records. As a result,
this will improve performance when you have multiple applications running
against a database and many commits are requested by the applications within a
very short time frame.
This grouping of commits will only occur when the value of this parameter is
greater than one and when the number of applications connected to the database is
greater than or equal to the value of this parameter. When commit grouping is
| This parameter should be incremented by small amounts only; for example one (1).
| You should also use multi-user tests to verify that increasing the value of this
| parameter provides the expected results.
Changes to the value specified for this parameter take effect immediately; you do
not have to wait until all applications disconnect from the database.
You could also sample the number of transactions per second and adjust this
parameter to accommodate the peak number of transactions per second (or some
large percentage of it). Accommodating peak activity would minimize the
overhead of writing log records during transaction intensive periods.
If you increase mincommit, you might also need to increase the logbufsz parameter
to avoid having a full log buffer force a write during these transaction intensive
periods. In this case, the logbufsz should be equal to:
mincommit * (log space used, on average, by a transaction)
You can use the database system monitor to help you tune this parameter in the
following ways:
v Calculating the peak number of transactions per second:
Taking monitor samples throughout a typical day, you can determine your
transaction intensive periods. You can calculate the total transactions by adding
the following monitor elements:
– commit_sql_stmts (commit statements attempted)
– rollback_sql_stmts (rollback statements attempted)
Using this information and the available timestamps, you can calculate the
number of transactions per second.
v Calculating the log space used per transaction:
Using sampling techniques over a period of time and a number of transactions,
you can calculate an average of the log space used with the following monitor
element:
– log_space_used (unit of work log space used)
Related reference:
v “uow_log_space_used - Unit of Work Log Space Used monitor element” in the
System Monitor Guide and Reference
v “commit_sql_stmts - Commit Statements Attempted monitor element” in the
System Monitor Guide and Reference
v “rollback_sql_stmts - Rollback Statements Attempted monitor element” in the
System Monitor Guide and Reference
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the number of times that DB2 is to try archiving a log file
| to the primary or the secondary archive directory before trying to archive log files
| to the failover directory. This parameter is only used if the failarchpath database
| configuration parameter is set. If numarchretry is not set, DB2 will continuously
| retry archiving to the primary or the secondary log path.
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
At the time of a database failure resulting from an event such as a power failure,
there might have been changes to the database which:
v Have not been committed, but updated the data in the buffer pool
v Have been committed, but have not been written from the buffer pool to the
disk
v Have been committed and written from the buffer pool to the disk.
When a database is restarted, the log files will be used to perform a crash recovery
of the database which ensures that the database is left in a consistent state (that is,
To determine which records from the log file need to be applied to the database,
the database manager uses a log control file. This log control file is periodically
written to disk, and, depending on the frequency of this event, the database
manager might be applying log records of committed transactions or applying log
records that describe changes that have already been written from the buffer pool
to disk. These log records have no impact on the database, but applying them
introduces some overhead into the database restart process.
The log control file is always written to disk when a log file is full, and during soft
checkpoints. You can use this configuration parameter to trigger additional soft
checkpoints.
The timing of soft checkpoints is based on the difference between the “current
state” and the “recorded state”, given as a percentage of the logfilsiz. The “recorded
state” is determined by the oldest valid log record indicated in the log control file
on disk, while the “current state” is determined by the log control information in
memory. (The oldest valid log record is the first log record that the recovery
process would read.) The soft checkpoint will be taken if the value calculated by
the following formula is greater than or equal to the value of this parameter:
( (space between recorded and current states) / logfilsiz ) * 100
Note however, that more page cleaner triggers and more frequent soft checkpoints
increase the overhead associated with database logging, which can impact the
performance of the database manager. Also, more frequent soft checkpoints might
not reduce the time required to restart a database, if you have:
v Very long transactions with few commit points.
v A very large buffer pool and the pages containing the committed transactions
are not written back to disk very frequently. (Note that the use of asynchronous
page cleaners can help avoid this situation.)
In both of these cases, the log control information kept in memory does not change
frequently and there is no advantage in writing the log control information to disk,
unless it has changed.
Related reference:
v “logfilsiz - Size of log files” on page 390
v “logprimary - Number of primary log files” on page 391
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
After logretain, or userexit, or both of these parameters are enabled, you must make
a full backup of the database. This state is indicated by the backup_pending flag
parameter.
Related reference:
v “User exit for database recovery” in the Data Recovery and High Availability Guide
and Reference
v “logretain - Log retain enable” on page 403
v “user_exit_status - User exit status indicator” on page 430
v “backup_pending - Backup pending indicator” on page 429
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “ROLLFORWARD DATABASE Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies additional parameters that DB2 might need to use to
| communicate with storage systems during backup, restore, or load copy
| operations.
| Related concepts:
| v Appendix E, “Cross-node recovery with the db2adutl command and the
| logarchopt1 and vendoropt database configuration parameters,” on page 589
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
Recovery
The following parameters affect various aspects of database recovery:
v “autorestart - Auto restart enable”
v “dft_loadrec_ses - Default number of load recovery sessions” on page 409
v “indexrec - Index re-creation time” on page 413
v “num_db_backups - Number of database backups” on page 415
v “rec_his_retentn - Recovery history retention period” on page 415
v “trackmod - Track modified pages enable” on page 416
The following parameters are used when working with Tivoli Storage Manager
(TSM):
v “tsm_mgmtclass - Tivoli Storage Manager management class” on page 416
v “tsm_nodename - Tivoli Storage Manager node name” on page 417
v “tsm_owner - Tivoli Storage Manager owner name” on page 417
v “tsm_password - Tivoli Storage Manager password” on page 418
When this parameter is set on, the database manager automatically calls the restart
database utility, if needed, when an application connects to a database. Crash
recovery is the operation performed by the restart database utility. It is performed if
the database terminated abnormally while applications were connected to it. An
abnormal termination of the database could be caused by a power failure or a
system software failure. It applies any committed transactions that were in the
database buffer pool but were not written to disk at the time of the failure. It also
backs out any uncommitted transactions that might have been written to disk.
Related concepts:
v “Crash recovery” in the Data Recovery and High Availability Guide and Reference
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “RESTART DATABASE Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies the default number of sessions that will be used during
the recovery of a table load. The value should be set to an optimal number of I/O
sessions to be used to retrieve a load copy. The retrieval of a load copy is an
operation similar to restore. You can override this parameter through entries in the
copy location file specified by the environment variable DB2LOADREC.
The default number of buffers used for load retrieval is two more than the value of
this parameter. You can also override the number of buffers in the copy location
file.
Related concepts:
v “Load Overview” in the Data Movement Utilities Guide and Reference
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “Miscellaneous variables” on page 518
| This parameter indicates the current role of a database, whether the database is
| online or offline. Valid values are: STANDARD, PRIMARY, or STANDBY.
| Note: Although the GET SNAPSHOT FOR DATABASE command returns high
| availability disaster recovery (HADR) status, it does so only when the
| database is online.
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the local host for high availability disaster recovery
| (HADR) TCP communication. Either a host name or an IP address can be used. If a
| host name is specified and it maps to multiple IP addresses, an error is returned,
| and HADR will not start up. If the host name maps to multiple IP addresses (even
| if you specify the same host name on primary and standby), primary and standby
| can end up mapping this host name to different IP addresses, because some DNS
| servers return IP address lists in non-deterministic order.
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the TCP/IP host name or IP address of the remote high
| availability disaster recovery (HADR) node. Similar to hadr_local_host, this
| parameter must map to only one IP address.
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the instance name of the remote server. Administration
| tools, such as the DB2 Control Center, use this parameter to contact the remote
| server. High availability disaster recovery (HADR) also checks whether a remote
| database requesting a connection belongs to the declared remote instance.
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the TCP service name or port number that will be used by
| the remote high availability disaster recovery (HADR) node.
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the synchronization mode, which determines how primary
| log writes are synchronized with the standby when the systems are in peer state.
| Valid values are:
| SYNC This mode provides the greatest protection against
| transaction loss, but at a higher cost of transaction
| response time.
| NEARSYNC This mode provides somewhat less protection
| against transaction loss, in exchange for a shorter
| transaction response time than that of SYNC mode.
| ASYNC This mode has the highest probability of
| transaction loss in the event of primary failure, in
| exchange for the shortest transaction response time
| among the three modes.
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| This parameter specifies the time (in seconds) that the high availability disaster
| recovery (HADR) process waits before considering a communication attempt to
| have failed.
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter indicates when the database manager will attempt to rebuild
invalid indexes. There are three possible settings for this parameter:
SYSTEM use system setting which will cause invalid indexes to be rebuilt at
the time specified in the database manager configuration file.
(Note: This setting is only valid for database configurations.)
ACCESS during index access which will cause invalid indexes to be rebuilt
when the index is first accessed.
Indexes can become invalid when fatal disk problems occur. If this happens to the
data itself, the data could be lost. However, if this happens to an index, the index
can be recovered by re-creating it. If an index is rebuilt while users are connected
to the database, two problems could occur:
v An unexpected degradation in response time might occur as the index file is
re-created. Users accessing the table and using this particular index would wait
while the index was being rebuilt.
v Unexpected locks might be held after index re-creation, especially if the user
transaction that caused the index to be re-created never performed a COMMIT
or ROLLBACK.
Recommendation: The best choice for this option on a high-user server and if
restart time is not a concern, would be to have the index rebuilt at DATABASE
RESTART time as part of the process of bringing the database back online after a
crash.
If this parameter is set to “RESTART”, the time taken to restart the database will be
longer due to index re-creation, but normal processing would not be impacted
once the database has been brought back online.
| Note: At database recovery time, all SQL procedure executables on the file system
| that belong to the database being recovered are removed. If indexrec is set to
| RESTART, all SQL procedure executables are extracted from the database
| catalog and put back on the file system at the next connection to the
| database. If indexrec is not set to RESTART, an SQL executable is extracted to
| the file system only on first execution of that SQL procedure.
Related tasks:
v “Backing up and restoring SQL procedures created prior to DB2 8.2” in the
Application Development Guide: Building and Running Applications
Related reference:
v “autorestart - Auto restart enable” on page 408
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “RESTART DATABASE Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the number of database backups to retain for a database.
After the specified number of backups is reached, old backups are marked as
expired in the recovery history file. Recovery history file entries for the table space
backups and load copy backups that are related to the expired database backup are
also marked as expired. When a backup is marked as expired, the physical
backups can be removed from where they are stored (for example, disk, tape,
TSM). The next database backup will prune the expired entries from the recovery
history file.
Related reference:
v “rec_his_retentn - Recovery history retention period” on page 415
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter is used to specify the number of days that historical information on
backups should be retained. If the recovery history file is not needed to keep track
of backups, restores, and loads, this parameter can be set to a small number.
If value of this parameter is -1, the recovery history file can only be pruned
explicitly using the available commands or APIs. If the value is not -1, the recovery
history file is pruned after every full database backup.
The the value of this parameter will override the value of the num_db_backups
parameter, but rec_his_retentn and num_db_backups must work together. If the value
for num_db_backups is large, the value for rec_his_retentn should be large enough to
support that number of backups.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “PRUNE HISTORY/LOGFILE Command” in the Command Reference
v “num_db_backups - Number of database backups” on page 415
When this parameter is set to ″Yes″, the database manager tracks database
modifications so that the backup utility can detect which subsets of the database
pages must be examined by an incremental backup and potentially included in the
backup image. After setting this parameter to ″Yes″, you must take a full database
backup in order to have a baseline against which incremental backups can be
taken. Also, if this parameter is enabled and if a table space is created, then a
backup must be taken which contains that table space. This backup could be either
a database backup or a table space backup. Following the backup, incremental
backups will be permitted to contain this table space.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
The Tivoli Storage Manager management class determines how the TSM server
should manage the backup versions of the objects being backed up.
When performing any TSM backup, before using the management class specified
by the database configuration parameter, TSM first attempts to bind the backup
object to the management class specified in the INCLUDE-EXCLUDE list found in
the TSM client options file. If a match is not found, the default TSM management
class specified on the TSM server will be used. TSM will then rebind the backup
object to the management class specified by the database configuration parameter.
Thus, the default management class, as well as the management class specified by
the database configuration parameter, must contain a backup copy group, or the
backup operation will fail.
This parameter is used to override the default setting for the node name associated
with the Tivoli Storage Manager (TSM) product. The node name is needed to allow
you to restore a database that was backed up to TSM from another node.
The default is that you can only restore a database from TSM on the same node
from which you did the backup. It is possible for the tsm_nodename to be
overridden during a backup done through DB2 (for example, with the BACKUP
DATABASE command).
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter is used to override the default setting for the owner associated
with the Tivoli Storage Manager (TSM) product. The owner name is needed to
allow you to restore a database that was backed up to TSM from another node. It
is possible for the tsm_owner to be overridden during a backup done through DB2
(for example, with the BACKUP DATABASE command).
The default is that you can only restore a database from TSM on the same node
from which you did the backup.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter is used to override the default setting for the password associated
with the Tivoli Storage Manager (TSM) product. The password is needed to allow
you to restore a database that was backed up to TSM from another node.
Note: If the tsm_nodename is overridden during a backup done with DB2 (for
example, with the BACKUP DATABASE command), the tsm_password might
also have to be set.
The default is that you can only restore a database from TSM on the same node
from which you did the backup. It is possible for the tsm_nodename to be
overridden during a backup done with DB2.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies the time interval in seconds for which a transaction
manager (TM), resource manager (RM) or sync point manager (SPM) should retry
the recovery of any outstanding indoubt transactions found in the TM, the RM, or
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter identifies the sync point manager (SPM) log file size in 4 KB pages.
The log file is contained in the spmlog sub-directory under sqllib and is created
the first time SPM is started.
Recommendation: The sync point manager log file size should be large enough to
maintain performance, but small enough to prevent wasted space. The size
required depends on the number of transactions using protected conversations,
and how often COMMIT or ROLLBACK is issued.
This parameter specifies the directory where the sync point manager (SPM) logs
are written. By default, the logs are written to the sqllib/spmlog directory, which,
in a high-volume transaction environment, can cause an I/O bottleneck. Use this
parameter to have the SPM log files placed on a faster disk than the current
sqllib/spmlog directory. This allows for better concurrency among the SPM agents.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter identifies the number of agents that can simultaneously perform
resync operations.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
This parameter identifies the name of the sync point manager (SPM) instance to
the database manager.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter identifies the name of the transaction manager (TM) database for
each DB2 instance. A TM database can be:
v A local DB2 Universal Database database
v A remote DB2 Universal Database database that does not reside on a host or
AS/400 system
v A DB2 for OS/390 Version 5 database if accessed via TCP/IP and the sync point
manager (SPM) is not used.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
Database management
A number of parameters provide information about your database or influence the
management of your database. These are grouped as follows:
v “Query Enabler”
v “Attributes” on page 423
v “DB2 Data Links Manager” on page 425
v “Status” on page 428
v “Compiler settings” on page 431
| v “Automated maintenance” on page 437
Query Enabler
The following parameter provides information for the control of Query Enabler:
v “dyn_query_mgmt - Dynamic SQL query management”
This parameter is relevant where DB2 Query Patroller is installed. If this parameter
is set to “ENABLE”, Query Patroller captures information about the query, such as
the submitter ID and the estimated cost of execution, as calculated by the
optimizer. These values are used to determine whether the query should be
managed by Query Patroller, based on user- and system-level thresholds.
If this parameter is set to “DISABLE”, Query Patroller does not capture any
information about submitted queries, and no query management takes place.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
With the exception of alt_collate, these parameters are provided for informational
purposes only.
This parameter specifies the collating sequence that is to be used for Unicode
tables in a non-Unicode database. Until this parameter is set, Unicode tables and
routines cannot be created in a non-Unicode database. Once set, this parameter
cannot be changed or reset.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter shows the code page that was used to create the database. The
codepage parameter is derived based on the codeset parameter.
Related reference:
v “codeset - Codeset for the database” on page 424
v “GET DATABASE CONFIGURATION Command” in the Command Reference
This parameter shows the codeset that was used to create the database. Codeset is
used by the database manager to determine codepage parameter values.
Related reference:
v “codepage - Code page for the database” on page 423
v “GET DATABASE CONFIGURATION Command” in the Command Reference
This parameter provides 260 bytes of database collating information. The first 256
bytes specify the database collating sequence, where byte “n” contains the sort
weight of the code point whose underlying decimal representation is “n” in the
code page of the database.
The last 4 bytes contain internal information about the type of the collating
sequence. You can treat it as an integer applicable to the platform of the database.
There are three values:
v 0 – The sequence contains non-unique weights
v 1 – The sequence contains all unique weights
v 2 – The sequence is the identity sequence, for which strings are compared byte
for byte.
If you use this internal type information, you need to consider byte reversal when
retrieving information for a database on a different platform.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
This parameter shows the territory code used to create the database.
Related reference:
v “territory - Database territory” on page 425
v “GET DATABASE CONFIGURATION Command” in the Command Reference
This parameter indicates the release level of the database manager which can use
the database. In the case of an incomplete or failed migration, this parameter will
reflect the release level of the unmigrated database and might differ from the
release parameter (the release level of the database configuration file). Otherwise
the value of database_level will be identical to value of the release parameter.
Related reference:
v “release - Configuration file release level” on page 425
v “GET DATABASE CONFIGURATION Command” in the Command Reference
Related reference:
v “database_level - Database release level” on page 424
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
This parameter shows the territory used to create the database. territory is used by
the database manager to determine the territory code (territory) parameter values.
Related reference:
v “country - Database territory code” on page 424
v “GET DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies whether Data Links support is enabled. A value of “YES”
specifies that Data Links support is enabled for Data Links Manager linking files
stored in native filesystems (for example, JFS on AIX). A value of “NO” specifies
that Data Links support is not enabled.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the interval of time (in seconds) for which the generated
file access control token is valid. The number of seconds the token is valid begins
from the time it is generated. The Data Links Filesystem Filter checks the validity
of the token against this expiry time.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies the number of additional copies of a file to be made in the
archive server (such as a TSM server) when a file is linked to the database.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies the interval of time (in days) files would be retained on an
archive server (such as a TSM server) after a DROP DATABASE is issued.
The default value for this parameter is one (1) day. A value of zero (0) means that
the files are deleted immediately from the archive server when the DROP
command is issued. (The actual file is not deleted unless the ON UNLINK
DELETE parameter was specified for the DATALINK column.)
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies the algorithm used in the generation of DATALINK file
access control tokens. The value of MAC1 (message authentication code) generates
a more secure message authentication code than MAC0, but also has more
performance overhead.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
The parameter indicates whether the file access control tokens use uppercase
letters. A value of “YES” specifies that all letters in an access control token are
uppercase. A value of “NO” specifies that the token can contain both uppercase
and lowercase letters.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies the initial expiry interval of a write token generated from
a DATALINK column. The initial expiry interval is the period between the time a
token is generated and the first time it is used to open the file.
Recommendation: The value should be large enough to cover the time from when
a write token is retrieved to when it is used to open the file.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
Status
The following parameters provide information about the state of the database:
v “backup_pending - Backup pending indicator” on page 429
v “database_consistent - Database is consistent” on page 429
v “log_retain_status - Log retain status indicator” on page 429
v “multipage_alloc - Multipage file allocation enabled” on page 429
v “restore_pending - Restore pending” on page 430
If set on, this parameter indicates that you must do a full backup of the database
before accessing it. This parameter is only on if the database configuration is
changed so that the database moves from being nonrecoverable to recoverable (that
is, initially both the logretain and userexit parameters were set to NO, then either
one or both of these parameters is set to YES, and the update to the database
configuration is accepted).
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
YES indicates that all transactions have been committed or rolled back so that the
data is consistent. If the system “crashes” while the database is consistent, you do
not need to take any special action to make the database usable.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
If set, this parameter indicates that log files are being retained for use in
roll-forward recovery.
This parameter is set when the logretain parameter setting is equal to Recovery.
Related reference:
v “logretain - Log retain enable” on page 403
v “GET DATABASE CONFIGURATION Command” in the Command Reference
| The default for the parameter is Yes: multipage file allocation is enabled.
| Following database creation, this parameter cannot be set to No. Multipage file
| allocation cannot be disabled once it has been enabled. If multipage file allocation
| is not desired, the DB2_NO_MPFA_FOR_NEW_DB DB2 registry variable must be
| set appropriately before the database is created. The db2empfa tool can be used to
| enable multipage file allocation for a database that currently has it disabled.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “Performance variables” on page 506
This parameter states whether a RESTORE PENDING status exists in the database.
Related reference:
v “userexit - User exit enable” on page 406
v “GET DATABASE CONFIGURATION Command” in the Command Reference
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
If set to Yes, this indicates that the database manager is enabled for roll-forward
recovery and that the user exit program will be used to archive and retrieve log
files when called by the database manager.
Related reference:
Compiler settings
The following parameters provide information to influence the compiler:
v “dft_degree - Default degree”
v “dft_mttb_types - Default maintained table types for optimization”
v “dft_queryopt - Default query optimization class” on page 432
v “dft_refresh_age - Default refresh age” on page 433
v “dft_sqlmathwarn - Continue upon arithmetic exceptions” on page 433
v “num_freqvalues - Number of frequent values retained” on page 434
v “num_quantiles - Number of quantiles for columns” on page 435
This parameter specifies the default value for the CURRENT DEGREE special
register and the DEGREE bind option.
Related reference:
v “max_querydegree - Maximum query degree of parallelism” on page 450
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies the default value for the CURRENT MAINTAINED
TABLE TYPES FOR OPTIMIZATION special register. The value of this register
determines what types of refresh deferred materialized query tables will be used
during query optimization.
Related reference:
v “CURRENT MAINTAINED TABLE TYPES FOR OPTIMIZATION special
register” in the SQL Reference, Volume 1
The query optimization class is used to direct the optimizer to use different
degrees of optimization when compiling SQL queries. This parameter provides
additional flexibility by setting the default query optimization class used when
neither the SET CURRENT QUERY OPTIMIZATION statement nor the
QUERYOPT option on the bind command are used.
Related reference:
v “SET CURRENT QUERY OPTIMIZATION statement” in the SQL Reference,
Volume 2
v “BIND Command” in the Command Reference
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
v “LIST DRDA INDOUBT TRANSACTIONS Command” in the Command Reference
This parameter has the default value used for the REFRESH AGE if the CURRENT
REFRESH AGE special register is not specified. This parameter specifies a time
stamp duration value with a data type of DECIMAL(20,6). This time duration
represents the maximum duration since a REFRESH TABLE statement has been
processed on a specific REFRESH DEFERRED materialized query table during
which that summary table can be used to optimize the processing of a query. If the
CURRENT REFRESH AGE has a value of 99999999999999 (ANY), and the QUERY
OPTIMIZATION class has a value of two, or five or more, REFRESH DEFERRED
materialized query tables are considered to optimize the processing of a dynamic
SQL query.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter sets the default value that determines the handling of arithmetic
errors and retrieval conversion errors as errors or warnings during SQL statement
compilation. For static SQL statements, the value of this parameter is associated
with the package at bind time. For dynamic SQL DML statements, the value of this
parameter is used when the statement is prepared.
Attention: If you change the dft_sqlmathwarn value for a database, the behavior of
check constraints, triggers, and views that include arithmetic expressions might
change. This might, in turn, have an impact on the data integrity of the database.
You should only change the setting of dft_sqlmathwarn for a database after carefully
evaluating how the new arithmetic exception handling behavior might impact
check constraints, triggers, and views. Once changed, subsequent changes require
the same careful evaluation.
When dft_sqlmathwarn is “No” and an INSERT with B=0 is attempted, the division
by zero is processed as an arithmetic error. The insert operation fails because DB2
cannot check the constraint. If dft_sqlmathwarn is changed to “Yes”, the division by
zero is processed as an arithmetic warning with a NULL result. The NULL result
causes the “>” predicate to evaluate to UNKNOWN and the insert operation
succeeds. If dft_sqlmathwarn is changed back to “No”, an attempt to insert the same
row will fail, because the division by zero error prevents DB2 from evaluating the
constraint. The row inserted with B=0 when dft_sqlmathwarn was “Yes” remains in
Before changing dft_sqlmathwarn from “Yes” to “No”, you should first check for
data that might become inconsistent by using, for example, predicates such as the
following:
WHERE A IS NOT NULL AND B IS NOT NULL AND A/B IS NULL
When inconsistent rows are isolated, you should take appropriate action to correct
the inconsistency before changing dft_sqlmathwarn. You can also manually re-check
constraints with arithmetic expressions after the change. To do this, first place the
affected tables in a check pending state (with the OFF clause of the SET
CONSTRAINTS statement), then request that the tables be checked (with the
IMMEDIATE CHECKED clause of the SET CONSTRAINTS statement). Inconsistent
data will be indicated by an arithmetic error, which prevents the constraint from
being evaluated.
Recommendation: Use the default setting of no, unless you specifically require
queries to be processed that include arithmetic exceptions. Then specify the value
of yes. This situation can occur if you are processing SQL statements that, on other
database managers, provide results regardless of the arithmetic exceptions that
occur.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter allows you to specify the number of “most frequent values” that
will be collected when the WITH DISTRIBUTION option is specified on the
RUNSTATS command. Increasing the value of this parameter increases the amount
of statistics heap (stat_heap_sz) used when collecting statistics.
You can also specify the number of frequent values retained as part of the
RUNSTATS command at the table or the column level. If none is specified, the
num_freqvalues configuration parameter value is used.
Updating this parameter can help the optimizer obtain better selectivity estimates
for some predicates (=, <, >, IS NULL, IS NOT NULL) over data that is
non-uniformly distributed. More accurate selectivity calculations might result in
the choice of more efficient access plans.
The RUNSTATS command allows for the specification of the number of frequent
values retained, by using the NUM_REQVALUES option. Changing the number of
frequent values retained through the RUNSTATS command is easier than making
the change using the num_freqvalues database configuration parameter.
When using RUNSTATS, you have the ability to limit the number of frequent
values collected at both the table level and the column level. This allows you to
optimize on space occupied in the catalogs by reducing the distribution statistics
for columns where they could not be exploited and yet still using the information
for critical columns.
Note that the process of collecting frequent value statistics requires significant CPU
and memory (stat_heap_sz) resources.
Related reference:
v “num_quantiles - Number of quantiles for columns” on page 435
v “stat_heap_sz - Statistics heap size” on page 356
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter controls the number of quantiles that will be collected when the
WITH DISTRIBUTION option is specified on the RUNSTATS command. Increasing
the value of this parameter increases the amount of statistics heap (stat_heap_sz)
used when collecting statistics.
The “quantile” statistics help the optimizer understand the distribution of data
values within a column. A higher value results in more information being available
to the SQL optimizer but requires additional catalog space. When 0 or 1 is
specified, no quantile statistics are retained, even if you request that distribution
statistics be collected.
You can also specify the number of quantiles collected as part of the RUNSTATS
command at the table or the column level. If none is specified, the num_quantiles
configuration parameter value is used.
Updating this parameter can help obtain better selectivity estimates for range
predicates over data that is non-uniformly distributed. Among other optimizer
decisions, this information has a strong influence on whether an index scan or a
table scan will be chosen. (It is more efficient to use a table scan to access a range
of values that occur frequently and it is more efficient to use an index scan for a
range of values that occur infrequently.)
The RUNSTATS command allows for the specification of the number of quantiles
that will be collected, by using the NUM_QUANTILES option. Changing the
number of quantiles that will be collected through the RUNSTATS command is
easier than making the change using the num_quantiles database configuration
parameter.
When using RUNSTATS, you have the ability to limit the number of quantiles
collected at both the table level and the column level. This allows you to optimize
on space occupied in the catalogs by reducing the distribution statistics for
columns where they could not be exploited and yet still using the information for
critical columns.
Related reference:
v “num_freqvalues - Number of frequent values retained” on page 434
v “stat_heap_sz - Statistics heap size” on page 356
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
| Automated maintenance
| The following parameter allows you to control the automatic maintenance activities
| of several DB2 utilities:
| v “autonomic_switches - Automatic maintenance switches”
| This parameter allows you to set a number of switches that are each internally
| represented by a bit of the parameter. You can update each of these switches
| independently by setting the following parameters:
| auto_maint This parameter is the parent of all the other
| automatic maintenance database configuration
| parameters (auto_db_backup, auto_tbl_maint,
| auto_runstats, auto_stats_prof, auto_prof_upd, and
| auto_reorg). When this parameter is disabled, all of
| its children parameters are also disabled, but their
| settings, as recorded in the database configuration
| file, do not change. When this parent parameter is
| enabled, recorded values for its children
| parameters take effect. In this way, automatic
| maintenance can be enabled or disabled globally.
| auto_db_backup This automated maintenance parameter enables or
| disables automatic backup operations for a
| database. A backup policy (a defined set of rules or
| guidelines) can be used to specify the automated
| behavior. The objective of the backup policy is to
| ensure that the database is being backed up
| regularly. The backup policy for a database is
| created automatically when the DB2 Health
| Related reference:
| v “GET DATABASE CONFIGURATION Command” in the Command Reference
| v “RESET DATABASE CONFIGURATION Command” in the Command Reference
| v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter allows you to assign a unique name to the database instance on a
workstation in the NetBIOS LAN environment. This nname is the basis for the
actual NetBIOS names that will be registered with NetBIOS for a workstation.
Since the NetBIOS protocol establishes connections using these NetBIOS names, the
nname parameter must be set for both the client and server.
Client applications must know the nname of the server that contains the database
to be accessed. The server’s nname must be cataloged in the client’s node directory
as the “server-nname” parameter using the CATALOG NETBIOS NODE command.
If nname at the server node changes to a new name, all clients accessing databases
on that server must catalog this new name for the server.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “CATALOG NETBIOS NODE Command” in the Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter contains the name of the TCP/IP port which a database server will
use to await communications from remote client nodes. This name must be the
first of two consecutive ports reserved for use by the database manager; the second
port is used to handle interrupt requests from down-level clients.
In order to accept connection requests from a database client using TCP/IP, the
database server must be listening on a port designated to that server. The system
administrator for the database server must reserve a port (number n) and define its
associated TCP/IP service name in the services file at the server. If the database
server needs to support requests from down-level clients, a second port (number
n+1, for interrupt requests) needs to be defined in the services file at the server.
The database server port (number n) and its TCP/IP service name need to be
defined in the services file on the database client. Down-level clients also require
the interrupt port (number n+1) to be defined in the client’s services file.
The svcename parameter should be set to the service name associated with the main
connection port so that when the database server is started, it can determine on
which port to listen for incoming connection requests. If you are supporting or
using a down-level client, the service name for the interrupt port is not saved in
the configuration file. The interrupt port number can be derived based on the main
connection port number (interrupt port number = main connection port + 1).
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter must be the same as the transaction program name that is
configured in the SNA transaction program definition.
Recommendation: The only accepted characters for use in this name are:
v Alphabetics (A through Z; or a through z)
v Numerics (0 through 9)
v Dollar sign ($), number sign (#), at sign (@), and period (.)
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
DB2 Discovery
You can use the following parameters to establish DB2 Discovery:
v “discover - Discovery mode”
v “discover_db - Discover database” on page 442
v “discover_inst - Discover server instance” on page 442
The default for this parameter is that discovery is enabled for this database.
Related reference:
v “GET DATABASE CONFIGURATION Command” in the Command Reference
v “RESET DATABASE CONFIGURATION Command” in the Command Reference
v “UPDATE DATABASE CONFIGURATION Command” in the Command Reference
This parameter specifies whether this instance can be detected by DB2 discovery.
The default, enable, specifies that the instance can be detected, while disable
prevents the instance from being discovered.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
Communications
The following parameters provide information about communications in the
partitioned database environment:
v “conn_elapse - Connection elapse time”
v “fcm_num_anchors - Number of FCM message anchors” on page 444
v “fcm_num_buffers - Number of FCM buffers” on page 444
v “fcm_num_connect - Number of FCM connection entries” on page 446
v “fcm_num_rqb - Number of FCM request blocks” on page 446
v “max_connretries - Node connection retries” on page 447
v “max_time_diff - Maximum time difference among nodes” on page 448
v “start_stop_time - Start and stop timeout” on page 448
This parameter specifies the number of seconds within which a TCP/IP connection
is to be established between two database partition servers. If the attempt
completes within the time specified by this parameter, communications are
established. If it fails, another attempt is made to establish communications. If the
connection is attempted the number of times specified by the max_connretries
parameter and always times out, an error is issued.
Related reference:
v “max_connretries - Node connection retries” on page 447
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the number of FCM message anchors. Agents use the
message anchors to send messages among themselves. The default value (-1)
specifies 75 percent of the value specified for fcm_num_rqb.
Related concepts:
v “Fast communications manager (FCM) communications” in the Administration
Guide: Implementation
Related reference:
v “fcm_num_buffers - Number of FCM buffers” on page 444
v “fcm_num_connect - Number of FCM connection entries” on page 446
v “fcm_num_rqb - Number of FCM request blocks” on page 446
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the number of 4 KB buffers that are used for internal
communications (messages) both among and within database servers.
If you have multiple logical nodes on the same machine, you might find it
necessary to increase the value of this parameter. You might also find it necessary
to increase the value of this parameter if you run out of message buffers because of
the number of users on the system, the number of database partition servers on the
system, or the complexity of the applications.
If you are using multiple logical nodes, on non-AIX systems, one pool of
fcm_num_buffers buffers is shared by all the multiple logical nodes on the same
machine, while on AIX:
v If there is enough room in the general memory that is used by the database
manager, the FCM buffer heap will be allocated from there. In this situation,
each database partition server will have fcm_num_buffers buffers of its own; the
database partition servers will not share a pool of FCM buffers (this was new in
DB2 Version 5).
v If there is not enough room in the general memory that is used by the database
manager, the FCM buffer heap will be allocated from a separate memory area
(AIX shared memory set), that is shared by all the multiple logical nodes on the
same machine. One pool of fcm_num_buffers will be shared by all the multiple
logical nodes on the same machine. This is the default configuration for all
non-AIX platforms.
Re-examine the value you are using; consider how many FCM buffers in total will
be allocated on the machine (or machines) where the multiple logical nodes reside.
Related concepts:
v “Fast communications manager (FCM) communications” in the Administration
Guide: Implementation
Related reference:
v “fcm_num_connect - Number of FCM connection entries” on page 446
v “fcm_num_anchors - Number of FCM message anchors” on page 444
v “fcm_num_rqb - Number of FCM request blocks” on page 446
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the number of FCM connection entries. Agents use
connection entries to pass data among themselves. The default value (-1) specifies
75 percent of the value specified for fcm_num_rqb.
Related concepts:
v “Fast communications manager (FCM) communications” in the Administration
Guide: Implementation
Related reference:
v “fcm_num_buffers - Number of FCM buffers” on page 444
v “fcm_num_anchors - Number of FCM message anchors” on page 444
v “fcm_num_rqb - Number of FCM request blocks” on page 446
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the number of FCM request blocks. Request blocks are the
media through which information is passed between the FCM daemon and an
agent, or between agents.
The requirement for request blocks will vary according to the number of users on
the system, the number of database partition servers in the system, and the
complexity of queries. Start with the default value, and use results from the
database system monitor when fine tuning this parameter.
Related concepts:
v “Fast communications manager (FCM) communications” in the Administration
Guide: Implementation
Related reference:
v “fcm_num_buffers - Number of FCM buffers” on page 444
v “fcm_num_connect - Number of FCM connection entries” on page 446
v “fcm_num_anchors - Number of FCM message anchors” on page 444
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
Related reference:
v “conn_elapse - Connection elapse time” on page 443
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
Each database partition server has its own system clock. This parameter specifies
the maximum time difference, in minutes, that is permitted among the database
partition servers listed in the node configuration file.
If two or more database partition servers are associated with a transaction, and
their clocks are not synchronized to within the time specified by this parameter,
the transaction is rejected and an SQLCODE is returned. (The transaction is
rejected only if data modification is associated with it.)
DB2 uses Coordinated Universal Time (UTC), so different time zones are not a
consideration when you set this parameter. The Coordinated Universal Time is the
same as Greenwich Mean Time.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the time, in minutes, within which all database partition
servers must respond to a DB2START or a DB2STOP command. It is also used as
the timeout value during an ADD DBPARTITIONNUM operation.
Database partition servers that do not respond to a DB2STOP command within the
specified time send a message to the db2stop error log in the log subdirectory of
the sqllib subdirectory of the home directory for the instance. You can either issue
DB2STOP for each database partition server that does not respond, or for all of
them. (Those that are already stopped will return stating that they are stopped.)
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “ADD DBPARTITIONNUM Command” in the Command Reference
Parallel processing
The following parameters provide information about parallel processing:
v “intra_parallel - Enable intra-partition parallelism”
v “max_querydegree - Maximum query degree of parallelism” on page 450
This parameter specifies whether the database manager can use intra-partition
parallelism.
Related reference:
v “max_querydegree - Maximum query degree of parallelism” on page 450
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
The default value for this configuration parameter is -1. This value means that the
system uses the degree of parallelism determined by the optimizer; otherwise, the
user-specified value is used.
Note: The degree of parallelism for an SQL statement can be specified at statement
compilation time using the CURRENT DEGREE special register or the
DEGREE bind option.
Related reference:
v “intra_parallel - Enable intra-partition parallelism” on page 449
v “dft_degree - Default degree” on page 431
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
Instance management
A number of parameters can help you manage your database manager instances.
These are grouped into the following categories:
v “Diagnostic”
v “Database system monitor parameters” on page 455
v “System management” on page 456
v “Instance administration” on page 464
Diagnostic
The following parameters allow you to control diagnostic information available
from the database manager:
v “diaglevel - Diagnostic error capture level”
v “diagpath - Diagnostic data directory path” on page 452
v “health_mon - Health monitoring” on page 453
v “notifylevel - Notify level” on page 453
This parameter specifies the type of diagnostic errors that will be recorded in the
[Link] file. Valid values are:
0 – No diagnostic data captured
1 – Severe errors only
2 – All errors
The diagpath configuration parameter is used to specify the directory that will
contain the error file, event log file (on Windows NT only), alert log file, and any
dump files that might be generated, based on the value of the diaglevel parameter.
Related reference:
v “diagpath - Diagnostic data directory path” on page 452
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter allows you to specify the fully qualified path for DB2 diagnostic
information. This directory could possibly contain dump files, trap files, an error
log, a notification file, and an alert log file, depending on your platform.
If this parameter is null, the diagnostic information will be written to files in one
of the following directories or folders:
v For supported Windows environments:
– If the DB2INSTPROF environment variable or keyword is not set, information
will be written to x:\SQLLIB\DB2INSTANCE, where x:\SQLLIB is the drive
reference and directory specified in the DB2PATH registry variable or
environment variable, and DB2INSTANCE is the name of the instance owner.
Related reference:
v “diaglevel - Diagnostic error capture level” on page 451
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter allows you to specify whether you want to monitor an instance, its
associated databases, and database objects according to various health indicators. If
health_mon is turned on, an agent will collect information about the health of the
objects you have selected. If an object is considered to be in an unhealthy position,
based on thresholds that you have set, notifications can be sent, and actions can be
taken automatically. If health_mon is turned off (the default), the health of objects
will not be monitored.
You can use the Health Center or the CLP to select the instance and database
objects that you want to monitor. You can also specify where notifications should
be sent, and what actions should be taken, based on the data collected by the
health monitor.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the type of administration notification messages that are
written to the administration notification log. On UNIX platforms, the
administration notification log is a text file called [Link]. On Windows, all
administration notification messages are written to the Event Log. The errors can
be written by DB2, the Health Monitor, the Capture and Apply programs, and user
applications.
For a user application to be able to write to the notification file or Windows Event
Log, it must call the db2AdminMsgWrite API.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter is unique in that it allows you to set a number of switches which
are each internally represented by a bit of the parameter. You can update each of
these switches independently by setting the following parameters:
dft_mon_uow Default value of the snapshot monitor’s unit of
work (UOW) switch
dft_mon_stmt Default value of the snapshot monitor’s statement
switch
dft_mon_table Default value of the snapshot monitor’s table
switch
dft_mon_bufpool Default value of the snapshot monitor’s buffer pool
switch
dft_mon_lock Default value of the snapshot monitor’s lock switch
dft_mon_sort Default value of the snapshot monitor’s sort switch
dft_mon_timestamp Default value of the snapshot monitor’s timestamp
switch
All monitoring applications inherit these default switch settings when the
application issues its first monitoring request (for example, setting a switch,
activating the event monitor, taking a snapshot). You should turn on a switch in
the configuration file only if you want to collect data starting from the moment the
database manager is started. (Otherwise, each monitoring application can set its
own switches and the data it collects becomes relative to the time its switches are
set.)
System management
The following parameters relate to system management:
v “comm_bandwidth - Communications bandwidth”
v “cpuspeed - CPU speed” on page 457
v “dft_account_str - Default charge-back account” on page 458
v “federated - Federated database system support” on page 458
v “jdk_path - Software Developer’s Kit for Java installation path” on page 459
v “nodetype - Machine node type” on page 459
v “numdb - Maximum number of concurrently active databases including host
and iSeries databases” on page 460
v “tp_mon_name - Transaction processor monitor name” on page 461
v “util_impact_lim - Instance impact policy” on page 463
The value calculated for the communications bandwidth, in megabytes per second,
is used by the SQL optimizer to estimate the cost of performing certain operations
between the database partition servers of a partitioned database system. The
optimizer does not model the cost of communications between a client and a
server, so this parameter should reflect only the nominal bandwidth between the
database partition servers, if any.
You can explicitly set this value to model a production environment on your test
system or to assess the impact of upgrading hardware.
Recommendation: You should only adjust this parameter if you want to model a
different environment.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
The CPU speed, in milliseconds per instruction, is used by the SQL optimizer to
estimate the cost of performing certain operations. The value of this parameter is
set automatically when you install the database manager based on the output from
a program designed to measure CPU speed. This program is executed, if
benchmark results are not available for any of the following reasons:
v The platform does not have support for the [Link] file
v The [Link] file is not found
v The data for the IBM RISC System/6000 model 530H is not found in the file
v The data for your machine is not found in the file.
You can explicitly set this value to model a production environment on your test
system or to assess the impact of upgrading hardware. By setting it to -1, cpuspeed
will be re-computed.
Recommendation: You should only adjust this parameter if you want to model a
different environment.
The CPU speed is used by the optimizer in determining access paths. You should
consider rebinding applications (using the REBIND PACKAGE command) after
changing this parameter.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
The suffix is supplied by the application program calling the sqlesact() API or
the user setting the environment variable DB2ACCOUNT. If a suffix is not
supplied by either the API or environment variable, DB2 Connect uses the value of
this parameter as the default suffix value. This parameter is particularly useful for
down-level database clients (anything prior to version 2) that do not have the
capability to forward an accounting string to DB2 Connect.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies the directory under which the Software Developer’s Kit
(SDK) for Java, to be used for running Java stored procedures and user-defined
functions, is installed. The CLASSPATH and other environment variables used by
the Java interpreter are computed from the value of this parameter.
Because there is no default value for this parameter, you should specify a value
when you install the SDK for Java.
Related reference:
v “java_heap_sz - Maximum Java interpreter heap size” on page 365
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter provides information about the DB2 products which you have
installed on your machine and, as a result, information about the type of database
manager configuration. The following are the possible values returned by this
parameter and the products associated with that node type:
v Database server with local and remote clients – a DB2 server product,
supporting local and remote database clients, and capable of accessing other
remote database servers.
v Client – a database client capable of accessing remote database servers.
v Database server with local clients – a DB2 relational database management
system, supporting local database clients and capable of accessing other, remote
database servers.
v Partitioned database server with local and remote clients – a DB2 server
product, supporting local and remote database clients, and capable of accessing
other remote database servers, and capable of partition parallelism.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
This parameter specifies the number of local databases that can be concurrently
active (that is, have applications connected to them), or the maximum number of
different database aliases that can be cataloged on a DB2 Connect server. Each
database takes up storage, and an active database uses a new shared memory
segment.
Related reference:
v “app_ctl_heap_sz - Application control heap size” on page 346
v “sortheap - Sort heap size” on page 355
v “aslheapsz - Application support layer heap size” on page 358
v “applheapsz - Application heap size” on page 350
v “locklist - Maximum storage for lock list” on page 340
v “dbheap - Database heap” on page 339
v “stmtheap - Statement heap size” on page 357
v “mon_heap_sz - Database system monitor heap size” on page 366
v “stat_heap_sz - Statistics heap size” on page 356
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “database_memory - Database shared memory size” on page 338
IBM WebSphere EJB and Microsoft Transaction Server users do not need to
configure any value for this parameter.
If none of the above products are being used, this parameter should not be
configured but left blank.
The maximum length of the string that can be specified for this parameter is 19
characters.
Related reference:
This parameter allows the database administrator (DBA) to limit the performance
degradation of a throttled utility on the workload. The DBA can then run online
utilities during critical production periods, and be guaranteed that the performance
impact on production work will be within acceptable limits.
A throttled utility will usually take longer to complete than an unthrottled utility.
If you find that a utility is running for an excessively long time, increase the value
of util_impact_lim, or disable throttling altogether by setting util_impact_lim to 100.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies and determines how and where authentication of a user
takes place.
If authentication is SERVER, the user ID and password are sent from the client to
the server so that authentication can take place on the server. The value
SERVER_ENCRYPT provides the same behavior as SERVER, except that any
passwords sent over the network are encrypted.
A value of CLIENT indicates that all authentication takes place at the client. No
authentication needs to be performed at the server.
Related reference:
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter specifies whether users are able to catalog and uncatalog databases
and nodes, or DCS and ODBC directories, without SYSADM authority. The default
value (0) for this parameter indicates that SYSADM authority is required. When
this parameter is set to 1 (yes), SYSADM authority is not required.
| This parameter specifies the name of the default Kerberos plug-in library to be
| used for client-side authentication and local authorization. By default, the value is
| null on UNIX-based systems, and IBMkrb5 on Windows operating systems. This
| plug-in is used when the client is authenticated using KERBEROS authentication.
| Related reference:
| v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
| Reference
| v “RESET DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
| v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
This parameter contains the default file path used to create databases under the
database manager. If no path is specified when a database is created, the database
is created under the path specified by the dftdbpath parameter.
In a partitioned database environment, you should ensure that the path on which
the database is being created is not an NFS-mounted path (on UNIX-based
platforms), or a network drive (in a Windows environment). The specified path
must physically exist on each database partition server. To avoid confusion, it is
best to specify a path that is locally mounted on each database partition server.
The maximum length of the path is 205 characters. The system appends the node
name to the end of the path.
Given that databases can grow to a large size and that many users could be
creating databases (depending on your environment and intentions), it is often
convenient to be able to have all databases created and stored in a specified
location. It is also good to be able to isolate databases from other applications and
data both for integrity reasons and for ease of backup and recovery.
For UNIX-based environments, the length of the dftdbpath name cannot exceed 215
characters and must be a valid, absolute, path name. For Windows, the dftdbpath
can be a drive letter, optionally followed by a colon.
Related reference:
Related reference:
v “authentication - Authentication type” on page 464
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “federated - Federated database system support” on page 458
| Related reference:
| v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
| Reference
| v “RESET DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
| v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
| This parameter specifies the name of the default GSS API plug-in library to be
| used for instance level local authorization when the value of the authentication
| database manager configuration parameter is set to GSSPLUGIN or
| GSS_SERVER_ENCRYPT.
| Related reference:
| v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
| Reference
| v “RESET DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
| v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
| Related reference:
| v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
| Reference
| v “RESET DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
| v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
| This parameter specifies the GSS API plug-in libraries that are supported by the
| database server. By default, the value is null. If the authentication type is
| GSSPLUGIN and this parameter is NULL, an error is returned. If the
| authentication type is KERBEROS and this parameter is NULL, the DB2-supplied
| kerberos module or library is used. This parameter is not used if another
| authentication type is used.
| When the authentication type is KERBEROS and the value of this parameter is not
| NULL, the list must contain exactly one Kerberos plug-in, and that plug-in is used
| for authentication (all other GSS plug-ins in the list are ignored). If there is more
| than one Kerberos plug-in, an error is returned.
| Each GSS API plug-in name must be separated by a comma (,) with no space
| either before or after the comma. Plug-in names should be listed in the order of
| preference. This parameter handles incoming connections at the server when the
| srvcon_auth parameter is specified as KERBEROS, KRB_SERVER_ENCRYPT,
| GSSPLUGIN or GSS_SERVER_ENCRYPT, or when srvcon_auth is not specified, and
| authentication is specified as KERBEROS, KRB_SERVER_ENCRYPT, GSSPLUGIN
| or GSS_SERVER_ENCRYPT.
| Related reference:
| v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
| Reference
| v “RESET DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
| This parameter specifies the name of the default userid-password plug-in library to
| be used for server-side authentication. By default, the value is null and the
| DB2-supplied userid-password plug-in library is used.
| The parameter handles incoming connections at the server when the srvcon_auth
| parameter is specified as SERVER or SERVER_ENCRYPT, or when srvcon_auth is
| not specified, and authentication is specified as CLIENT, SERVER, or
| SERVER_ENCRYPT.
| Related reference:
| v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
| Reference
| v “RESET DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
| v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
| This parameter specifies whether plug-ins are to run in fenced mode or unfenced
| mode. Unfenced mode is the only supported mode.
| Related reference:
| v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
| Reference
| v “RESET DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
To restore the parameter to its default (NULL) value, use UPDATE DBM CFG
USING SYSADM_GROUP NULL. You must specify the keyword “NULL” in
uppercase. You can also use the Configure Instance notebook in the DB2 Control
Center.
Related reference:
v “sysctrl_group - System control authority group name” on page 473
v “sysmaint_group - System maintenance authority group name” on page 473
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter defines the group name with system control (SYSCTRL) authority.
SYSCTRL has privileges allowing operations affecting system resources, but does
not allow direct access to data.
Attention: This parameter must be NULL for Windows 98 clients when system
security is used (that is, authentication is CLIENT, SERVER, DCS, or any other
valid authentication). This is because the Windows 98 operating systems do not
store group information, thereby providing no way of determining if a user is a
member of a designated SYSCTRL group. When a group name is specified, no user
can be a member of it.
To restore the parameter to its default (NULL) value, use UPDATE DBM CFG
USING SYSCTRL_GROUP NULL. You must specify the keyword “NULL” in
uppercase. You can also use the Configure Instance notebook in the DB2 Control
Center.
Related reference:
v “sysadm_group - System administration authority group name” on page 472
v “sysmaint_group - System maintenance authority group name” on page 473
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
Attention: This parameter must be NULL for Windows 98 clients when system
security is used (that is, authentication is CLIENT, SERVER, DCS, or any other
valid authentication). This is because the Windows 98 operating systems do not
store group information, thereby providing no way of determining if a user is a
member of a designated SYSMAINT group. When a group name is specified, no
user can be a member of it.
To restore the parameter to its default (NULL) value, use UPDATE DBM CFG
USING SYSMAINT_GROUP NULL. You must specify the keyword “NULL” in
uppercase. You can also use the Configure Instance notebook in the DB2 Control
Center.
Related reference:
v “sysadm_group - System administration authority group name” on page 472
v “sysctrl_group - System control authority group name” on page 473
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
| This parameter defines the group name with system monitor (SYSMON) authority.
| Users having SYSMON authority at the instance level have the ability to take
| database system monitor snapshots of a database manager instance or its
| databases. SYSMON authority includes the ability to use the following commands:
| v GET DATABASE MANAGER MONITOR SWITCHES
| v GET MONITOR SWITCHES
| v GET SNAPSHOT
| v LIST ACTIVE DATABASES
| v LIST APPLICATIONS
| v LIST DCS APPLICATIONS
| v RESET MONITOR
| v UPDATE MONITOR SWITCHES
| Attention: This parameter must be NULL for Windows 98 clients when system
| security is used (that is, authentication is CLIENT, SERVER, DCS, or any other
| valid authentication). This is because the Windows 98 operating systems do not
| store group information, thereby providing no way of determining if a user is a
| member of a designated SYSMON group. When a group name is specified, no user
| can be a member of it.
| To restore the parameter to its default (NULL) value, use UPDATE DBM CFG
| USING SYSMON_GROUP NULL. You must specify the keyword “NULL” in
| uppercase. You can also use the Configure Instance notebook in the DB2 Control
| Center.
| Related reference:
| v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
| Reference
| v “RESET DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
| v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
| Command Reference
This parameter is only active when the authentication parameter is set to CLIENT.
This parameter and trust_clntauth are used to determine where users are validated
to the database environment.
By accepting the default of “YES” for this parameter, all clients are treated as
trusted clients. This means that the server assumes that a level of security is
available at the client and the possibility that users can be validated at the client.
This parameter can only be changed to “NO” if the authentication parameter is set
to CLIENT. If this parameter is set to “NO”, the untrusted clients must provide a
userid and password combination when they connect to the server. Untrusted
clients are operating system platforms that do not have a security subsystem for
authenticating users.
Setting this parameter to “DRDAONLY” protects against all clients except clients
from DB2 for OS/390 and z/OS, DB2 for VM and VSE, and DB2 for OS/400. Only
these clients can be trusted to perform client-side authentication. All other clients
must provide a user ID and password to be authenticated by the server.
Related reference:
v “authentication - Authentication type” on page 464
v “trust_clntauth - Trusted clients authentication” on page 476
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
If this parameter is set to CLIENT (the default), the trusted client can connect
without providing a user ID and password combination, and the assumption is
that the operating system has already authenticated the user. If it is set to SERVER,
the user ID and password will be validated at the server.
The numeric value for CLIENT is 0. The numeric value for SERVER is 1.
Related reference:
v “authentication - Authentication type” on page 464
v “trust_allclnts - Trust all clients” on page 475
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
Related reference:
v “authentication - Authentication type” on page 464
v “GET DATABASE MANAGER CONFIGURATION Command” in the Command
Reference
v “RESET DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
v “UPDATE DATABASE MANAGER CONFIGURATION Command” in the
Command Reference
This parameter determines how and where authentication of a user takes place.
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
This parameter specifies the location where the contact information used for
notification by the Scheduler and the Health Monitor is stored. The location is
defined to be a DB2 administration server’s TCP/IP hostname. Allowing
contact_host to be located on a remote DAS provides support for sharing a contact
list across multiple DB2 administration servers. If contact_host is not specified, the
DAS assumes the contact information is local.
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
This parameter indicates the code page used by the DB2 administration server. If
the parameter is null, then the default code page of the system is used. This
parameter should be compatible with the locale of the local DB2 instances.
Otherwise, the DB2 administration server cannot communicate with the DB2
instances.
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
v “das_territory - DAS territory” on page 479
This parameter shows the territory used by the DB2 administration server. If the
parameter is null, then the default territory of the system is used.
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
v “das_codepage - DAS code page” on page 479
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
This parameter specifies the name that is used by your users and database
administrators to identify the DB2 server system. If possible, this name should be
unique within your network.
This name is displayed in the system level of the Control Center’s object tree to aid
administrators in the identification of server systems that can be administered from
the Control Center.
When using the ’Search the Network’ function of the Configuration Assistant, DB2
discovery returns this name and it is displayed at the system level in the resulting
object tree. This name aids users in identifying the system that contains the
database they wish to access. A value for db2system is set at installation time as
follows:
v On Windows, the setup program sets it equal to the computer name specified
for the Windows system.
v On UNIX systems, it is set equal to the UNIX system’s TCP/IP hostname.
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
v “db2system - Name of the DB2 server system” on page 480
This parameter specifies whether or not the Scheduler will execute tasks that have
been scheduled in the past, but have not yet been executed. The Scheduler only
detects expired tasks when it starts up.
For example, if you have a job scheduled to run every Saturday, and the Scheduler
is turned off on Friday and then restarted on Monday, the job scheduled for
Saturday is now a job that is scheduled in the past. If exec_exp_task is set to Yes,
your Saturday job will run when the Scheduler is restarted.
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
v “sched_enable - Scheduler mode” on page 483
v “toolscat_inst - Tools catalog database instance” on page 485
v “toolscat_db - Tools catalog database” on page 485
v “toolscat_schema - Tools catalog database schema” on page 486
v “smtp_server - SMTP server” on page 484
v “sched_userid - Scheduler user ID” on page 484
This parameter specifies the directory under which the 64-Bit Software Developer’s
Kit (SDK) for Java, to be used for running DB2 administration server functions, is
installed.
Note: This is different from the jdk_path configuration parameter, which specifies a
32-bit SDK for Java.
Environment variables used by the Java interpreter are computed from the value of
this parameter. This parameter is only used on those platforms that support both
32- and 64-bit instances. Those platforms are also known as 64-bit hybrid
platforms, and include AIX, HP-UX, and the Solaris Operating Environment. On all
other platforms, only jdk_path is used.
Because there is no default value for this parameter, you should specify a value
when you install the SDK for Java.
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
This parameter specifies the directory under which the Software Developer’s Kit
(SDK) for Java, to be used for running DB2 administration server functions, is
installed. Environment variables used by the Java interpreter are computed from
the value of this parameter.
On Windows operating systems, Java files (if needed) are placed under the sqllib
directory (in java\jdk) during DB2 installation. The jdk_path configuration
parameter is then set to sqllib\java\jdk. Java is never actually installed by DB2 on
Windows platforms; the files are merely placed under the sqllib directory, and this
is done regardless of whether or not Java is already installed.
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
v “toolscat_inst - Tools catalog database instance” on page 485
v “toolscat_db - Tools catalog database” on page 485
v “toolscat_schema - Tools catalog database schema” on page 486
v “smtp_server - SMTP server” on page 484
v “exec_exp_task - Execute expired tasks” on page 481
v “sched_userid - Scheduler user ID” on page 484
This parameter specifies the user ID used by the Scheduler to connect to the tools
catalog database. This parameter is only relevant if the tools catalog database is
remote to the DB2 administration server.
The userid and password used by the Scheduler to connect to the remote tools
catalog database are specified using the db2admin command.
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “sched_enable - Scheduler mode” on page 483
v “toolscat_inst - Tools catalog database instance” on page 485
v “toolscat_db - Tools catalog database” on page 485
v “toolscat_schema - Tools catalog database schema” on page 486
v “smtp_server - SMTP server” on page 484
v “exec_exp_task - Execute expired tasks” on page 481
When the Scheduler is on, this parameter identifies the SMTP server that the
Scheduler will use to send e-mail and pager notifications.
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
v “sched_enable - Scheduler mode” on page 483
v “toolscat_inst - Tools catalog database instance” on page 485
v “toolscat_db - Tools catalog database” on page 485
v “toolscat_schema - Tools catalog database schema” on page 486
v “exec_exp_task - Execute expired tasks” on page 481
v “sched_userid - Scheduler user ID” on page 484
This parameter indicates the tools catalog database used by the Scheduler. This
database must be in the database directory of the instance specified by toolscat_inst.
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
v “sched_enable - Scheduler mode” on page 483
v “toolscat_inst - Tools catalog database instance” on page 485
v “toolscat_schema - Tools catalog database schema” on page 486
v “smtp_server - SMTP server” on page 484
v “exec_exp_task - Execute expired tasks” on page 481
v “sched_userid - Scheduler user ID” on page 484
This parameter indicates the instance name that is used by the Scheduler, along
with toolscat_db and toolscat_schema, to identify the tools catalog database. The
tools catalog database contains task information created by the Task Center and the
Control Center. The tools catalog database must be listed in the database directory
of the instance specified by this configuration parameter. The database can be local
or remote. If the tools catalog database is local, the instance must be configured for
TCP/IP. If the database is remote, the node cataloged in the database directory
must be a TCP/IP node.
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related tasks:
v “Tools catalog database and DAS scheduler setup and configuration” in the
Administration Guide: Implementation
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
This parameter indicates the schema of the tools catalog database used by the
Scheduler. The schema is used to uniquely identify a set of tools catalog tables and
views within the database.
This parameter can only be updated from a Version 8 command line processor
(CLP).
Related reference:
v “GET ADMIN CONFIGURATION Command” in the Command Reference
v “RESET ADMIN CONFIGURATION Command” in the Command Reference
v “UPDATE ADMIN CONFIGURATION Command” in the Command Reference
v “sched_enable - Scheduler mode” on page 483
v “toolscat_inst - Tools catalog database instance” on page 485
v “toolscat_db - Tools catalog database” on page 485
v “smtp_server - SMTP server” on page 484
v “exec_exp_task - Execute expired tasks” on page 481
v “sched_userid - Scheduler user ID” on page 484
To view a list of all supported registry variables, execute the following command:
db2set -lr
To change the value for a variable in the current or default instance, execute the
following command:
db2set registry_variable_name=new_value
You must set the values for the changed registry variables before you execute the
DB2START command.
Note: If a registry variable requires Boolean values as arguments, the values YES,
1, and ON are all equivalent and the the values NO, 0, and OFF are also
equivalent. For any variable, you can specify any of the appropriate
equivalent values.
Related concepts:
v “Environment variables and the profile registry” in the Administration Guide:
Implementation
Related tasks:
v “Setting DB2 registry variables at the user level in the LDAP environment” in
the Administration Guide: Implementation
Related reference:
v “General registry variables” on page 490
v “System environment variables” on page 492
v “Communications variables” on page 496
v “Command-line variables” on page 499
v “MPP configuration variables” on page 500
v “SQL compiler variables” on page 502
v “Performance variables” on page 506
Values: YES or NO
This variable enables bidirectional support and the DB2CODEPAGE variable is used to declare the code page to be
used. Refer to the National Language Support appendix for additional information on bidirectional support.
DB2CODEPAGE All Default: derived from the language ID, as
specified by the operating system.
Specifies the code page of the data presented to DB2 for database client application. The user should not set
DB2CODEPAGE unless explicitly stated in DB2 documents, or asked to do so by DB2 service. Setting
DB2CODEPAGE to a value not supported by the operating system can produce unexpected results. Normally, you
do not need to set DB2CODEPAGE because DB2 automatically derives the code page information from the
operating system.
DB2_COLLECT_TS_REC_INFO All Default=OFF
Values: YES or NO
This variable specifies whether DB2 will process all log files when rolling forward a table space, regardless of
whether the log files contain log records that affect the table space. To skip the log files known not to contain any
log records affecting the table space, set this variable to ″ON″.
DB2_COLLECT_TS_REC_INFO must be set before the log files are created and used so that the information
required for skipping log files is collected.
DB2CONSOLECP Windows Default= null
If FALSE, the current statement fails. The application can still commit the work completed by previous statements in
the unit of work, or it can roll back the work completed to undo the unit of work.
DB2GRAPHICUNICODESERVER All Default=OFF
Values: ON or OFF
This registry variable is used to accommodate existing applications written to insert graphic data into a Unicode
database. Its use is only needed for applications that specifically send sqldbchar (graphic) data in Unicode instead
of the code page of the client. (sqldbchar is a supported SQL data type in C and C++ that can hold a single
double-byte character.) When set to “ON”, you are telling the database that graphic data is coming in Unicode, and
the application expects to receive graphic data in Unicode.
DB2INCLUDE All Default=current directory
Specifies a path to be used during the processing of the SQL INCLUDE text-file statement during DB2 PREP
processing. It provides a list of directories where the INCLUDE file might be found. Refer to the Application
Development Guide for descriptions of how DB2INCLUDE is used in the different precompiled languages.
DB2INSTDEF Windows operating Default=DB2
systems
Sets the value to be used if DB2INSTANCE is not defined.
DB2INSTOWNER Windows NT Default=null
The registry variable created in the DB2 profile registry when the instance is first created. This variable is set to the
name of the instance-owning machine.
DB2_LIC_STAT_SIZE All Default=null
Range: 0 to 32 767
The registry variable determines the maximum size (in MBs) of the file containing the license statistics for the
system. A value of zero turns the license statistic gathering off. If the variable is not recognized or not defined, the
variable defaults to unlimited. The statistics are displayed using the License Center.
DB2LOCALE All Default: NO
Values: YES or NO
Specifies whether the default ″C″ locale of a process is restored to the default ″C″ locale after calling DB2 and
whether to restore the process locale back to the original ’C’ after calling a DB2 function. If the original locale was
not ’C’, then this registry variable is ignored.
Minimum=16 buffers
This variable is used for NetBIOS search discovery. The variable specifies the number of concurrent discovery
responses that can be received by a client. If the client receives more concurrent responses than are specified by this
variable, then the excess responses are discarded by the NetBIOS layer. The default is sixteen (16) NetBIOS receive
buffers. If a number less than the default value is chosen, then the default is used.
| DB2_OBJECT_TABLE_ENTRIES All Default=0
Values: 0–50000
| Specifies the expected number of objects in a table space. If you know that a large number of objects (for example,
| 1000 or more) will be created in a DMS table space, you should set this registry variable to the approximate number
| before creating the table space. This will reserve contiguous storage for object metadata during table space creation.
| Reserving contiguous storage reduces the chance that an online backup will block operations which update entries
| in the metadata (for example, CREATE INDEX, IMPORT REPLACE). It will also make resizing the table space easier
| because the metadata will be stored at the start of the table space.
| If the initial size of the table space is not large enough to reserve the contiguous storage, the table space creation
| will continue without the additional space reserved.
DB2OPTIONS All Default=null
Sets command line processor options.
DB2TERRITORY All Default: derived from the language ID, as
specified by the operating system.
Specifies the region, or territory code of the client application, which influences date and time formats.
| DB2_VIEW_REOPT_VALUES All Default=NO
Values: YES, NO
| This variable enables all users to store the cached values of a reoptimized SQL statement in the
| EXPLAIN_PREDICATE table when the statement is explained. When this variable is set to NO, only DBADM is
| allowed to save these values in the EXPLAIN_PREDICATE table.
Related concepts:
v “DB2 registry and environment variables” on page 489
Values: YES or NO
When you set this variable to NO, local DB2 Connect clients on a DB2 Connect Enterprise Edition machine are
forced to run within an agent. Some advantages of running within an agent are that local clients can be monitored
and that they can use SYSPLEX support.
DB2DOMAINLIST Windows NT server Default=null
only
Values: A list of Windows NT domain names
separated by commas (“,”).
This variable is effective only when CLIENT authentication is set in the Database Manager configuration and is
needed if a single signon from a Windows NT desktop is required in a Windows NT domain environment.
This registry variable should only be used under a pure Windows NT domain environment with DB2 servers and
clients running DB2 Universal Database Version 7.1 (or later).
DB2ENVLIST UNIX Default: null
Lists specific variable names for either stored procedures or user-defined functions. By default, the db2start
command filters out all user environment variables except those prefixed with DB2 or db2. If specific environment
variables must be passed to either stored procedures or user-defined functions, you can list the variable names in
the DB2ENVLIST environment variable. Separate each variable name by one or more spaces.
DB2INSTANCE All Default=DB2INSTDEF on Windows 32-bit
operating systems.
The environment variable used to specify the instance that is active by default. On UNIX, users must specify a
value for DB2INSTANCE.
DB2INSTPROF Windows operating Default: null
systems
The environment variable used to specify the location of the instance directory on Windows operating systems, if
different than DB2PATH.
DB2LIBPATH UNIX Default: null
DB2 constructs its own shared library path. If you wish to add a PATH into the engine’s library path (for example,
on AIX, a user-defined function requires a specific entry in LIBPATH) then you must set DB2LIBPATH. The actual
valueof DB2LIBPATH is appended to the end of the DB2-constructed shared library path.
DB2NODE All Default: null
Values: 1 to 999
Used to specify the target logical node of a DB2 Enterprise Server Edition database partition server that you want to
attach to or connect to. If this variable is not set, the target logical node defaults to the logical node which is defined
with port 0 on the machine.
| Values:
| v *[:number-of-disks] - meaning every table
| space will have parallel I/O enabled. A
| value or symbol can be provided after a
| colon to define the default number of
| disks per container for all tablespaces. A
| default of 6 (the value for a RAID-5
| device) will be used if the number of disks
| is not specified.
| v TablespaceID[:number-of-disks],... - a
| comma-separated list of defined table
| spaces. To define the number of disks per
| container for that table space, specify a
| value following a colon after each table
| space ID. This value can be a numeric
| value or one of the symbols described in
| the table below.
| v A combination of the above two value
| types, with each value separated by a
| comma
| This registry variable is used to change the way DB2 calculates the I/O parallelism of a tablespace. When I/O
| parallelism is enabled (either implicitly, by the use of multiple containers, or explicitly, by setting
| DB2_PARALLEL_IO), it is achieved by issuing the correct number of prefetch requests. Each prefetch request is an
| request for an extent of pages.
| If this registry variable is not set, the degree of parallelism of any table space will be the number of containers of
| the table space. For example, if DB2_PARALLEL_IO is set to null and a table space has four containers, then there
| will be four extent-sized prefetch requests issued.
| If this registry variable is set, then the degree of parallelism of the table space will be the ratio between the prefetch
| size and the extent size of this table space. For example, if DB2_PARALLEL_IO is set for a table space that has a
| prefetch size of 160 and an extent size of 32 pages, then there will be five extent-sized prefetch requests issued. A
| wildcard ″*″ character can be used to tell DB2 to calculate the I/O parallelism for all table spaces this way.
| In I/O subsystems that support stripping of physical spindles beneath each DB2 container (for example, with a
| RAID device), the number of disks underneath each DB2 container should be taken into account when choosing a
| prefetch size for the table space. The prefetch size should be calculated based on the following equation:
| Prefetch size = (number of containers) * (number of disks per container) * extent size
| DB2 automatically calculates the prefetch size of a table space using the above equation if the prefetch size of the
| table space is AUTOMATIC.
| The DB2_PARALLEL_IO registry variable can be used to tell DB2 the number of disks per container. For example, if
| DB2_PARALLEL_IO=″1:4″ and table space 1 has three containers, the extent size 32, and prefetch size AUTOMATIC,
| then the prefetch size will be calculated as 3 * 4 * 32 = 384 pages. The I/O parallelism of this table space will be 384
| / 32 which is 12. If the prefetch size of a table space is not AUTOMATIC this information about the number of
| disks per container will not be used.
| Any table space that is specified under DB2_PARALLEL_IO will by default be assumed to be using six as the
| number of disks per container, unless otherwise specified in the registry variable. For example, if
| DB2_PARALLEL_IO=*,1:3, then all table spaces will use six as the number of disks per container, except for table
| space 1, which will use three.
When specified, DB2 issues the SetProcessAffinityMask() api. If unspecified, the db2syscs process is associated with
all processors on the machine.
DB2_USE_PAGE_CONTAINER_TAG All Default: null
However, if you set this registry variable to ON when you use RAID devices for containers, I/O performance might
degrade. Because for RAID devices you create table spaces with an extent size equal to or a multiple of the RAID
stripe size, setting the DB2_USE_PAGE_CONTAINER_TAG to ON causes the extents not to line up with the RAID
stripes. As a result, an I/O request might need to access more physical disks than would be optimal. Users are
strongly advised against enabling this registry variable.
To activate changes to this registry variable, issue a DB2STOP command and then enter a DB2START command.
DB2_CLPHISTSIZE All Default: 20
Note: This registry variable is not set to the
default value during installation. Instead, the
code that makes use of this variable uses a
default value of 20 if the registry variable is
not set or if it is set to a value outside of the
valid [Link]: 1–500 inclusive
This variable determines the number of commands stored in the command history during CLP interactive sessions.
Because the command history is held in memory, a very high value for this variable might result in a performance
impact depending on the number and length of commands run in a session.
DB2_CLP_EDITOR All Default:
UNIX: ’vi’
Note: This registry variable is not set to the
default value during [Link], the
code that makes use of this variable uses a
default value if the registry variable is not
[Link]: Any valid editor that is located in
the operating system path
This variable determines the editor to be used when executing the EDIT command. From a CLP interactive session,
the EDIT command launches an editor pre-loaded with a user-specified command which can then be edited and
run.
Communications variables
Table 48. Communications Variables
Variable Name Operating System Values
Description
| DB2CHECKCLIENTINTERVAL All, server only Default=50
| Lower values represent more frequent checks. As a guide, for low frequency, use 100; for medium frequency, use 50;
| for high frequency use 10. Checking more frequently for client status while executing a database request lengthens
| the time taken to complete queries. If the DB2 workload is heavy (that is, it involves many internal requests), then
| setting DB2CHECKCLIENTINTERVAL to a low value has a greater impact on performance than in a situation
| where the workload is light and DB2 is waiting most of the time.
| In DB2 UDB, Version 8.1.4, the default value for DB2CHECKCLIENTINTERVAL is 50. Prior to version 8.1.4, the
| default value was 0.
DB2COMM All, server only Default=null
Values: 1-65535
This can be set to override the default port number of the db2jd. If this registry variable is set, the db2jd will
attempt to use it as the listening port of the db2jd when a db2jstrt is issued with no parameters. If this registry
variable is not set, and no port parameter is provided, the db2jd will start on the default port of 6789.
DB2NBADAPTERS Windows Default=0
Range: 0-15,
Values: 1-720
Lower values ensure that the NetBIOS protocol checkup runs more often, freeing up memory and other system
resources left when unexpected agent/session termination occurs.
DB2NBINTRLISTENS Windows server only Default=1
Values: 1-10
Setting DB2NBINTRLISTENS to a lower value conserves NetBIOS sessions and NCBs at the server. However, in an
environment where client interrupts are common, you may need to set DB2NBINTRLISTENS to a higher value in
order to be responsive to interrupting clients.
Note: Values specified are position sensitive; they relate to the corresponding value positions for
DB2NBADAPTERS.
DB2NBRECVBUFFSIZE Windows server only Default=4096 bytes
Range: 4096-65536
Specifies the size of the DB2 NetBIOS protocol receive buffers. These buffers are assigned to the NetBIOS receive
NCBs. Lower values conserve server memory, while higher values may be required when client data transfers are
larger.
DB2NBBRECVNCBS Windows server only Default=10
Range: 1-99
Specifies the number of NetBIOS ″receive_any″ commands (NCBs) that the server issues and maintains during
operation. This value is adjusted depending on the number of remote clients to which your server is connected.
Lower values conserve server resources.
Note: Each adapter in use can have its own unique receive NCB value specified by DB2NBBRECVNCBS. The
values specified are position sensitive; they relate to the corresponding value positions for DB2NBADAPTERS.
DB2NBRESOURCES Windows server only Default=null
Specifies the number of NetBIOS resources to allocate for DB2 use in a multi-context environment. This variable is
restricted to multi-context client operation.
DB2NBSENDNCBS Windows server only Default=6
Range: 1-720
Specifies the number of send NetBIOS commands (NCBs) that the server reserves for use. This value can be
adjusted depending on the number of remote clients your server is connected to. Setting DB2NBSENDNCBS to a
lower value will conserve server resources. However, you might need to set it to a higher value to prevent the
server from waiting to send to a remote client when all other send commands are in use.
DB2NBSESSIONS Windows server only Default=null
Range: 5-254
Specifies the number of sessions that DB2 should request to be reserved for DB2 use. The value of DB2NBSESSIONS
can be set to request a specific session for each adapter specified using DB2NBADAPTERS.
Note: Values specified are position sensitive; they relate to the corresponding value positions for
DB2NBADAPTERS.
Range: 5-254
Specifies the number of ″extra″ NetBIOS commands (NCBs) the server will need to reserve when the db2start
command is issued. The value of DB2NBXTRANCBS can be set to request a specific session for each adapter
specified using DB2NBADAPTERS.
DB2RETRY Windows Default=0
Values: 1 to 8
Values: ON or OFF
Specifies whether to use the Virtual Interface (VI) Architecture communication protocol or not. If this registry
variable is “ON”, then FCM will use VI for inter-node communication. If this registry variable is “OFF”, then FCM
will use TCP/IP for inter-node communication.
Note: The value of this registry variable must be the same across all the database partitions in the instance.
DB2_VI_VIPL Windows NT Default= [Link]
Specifies the name of the Virtual Interface Provider Library (VIPL) that will be used by DB2. In order to load the
library successfully, the library name used in this registry variable must be in the PATH user environment variable.
The currently supported implementations all use the same library name.
DB2_VI_DEVICE Windows NT Default=null
Related concepts:
v “DB2 registry and environment variables” on page 489
Command-line variables
Table 49. Command Line Variables
Variable Name Operating System Values
Description
DB2BQTIME All Default=1 second
Related concepts:
v “DB2 registry and environment variables” on page 489
This registry variable is no longer needed, but is retained for backward compatibility.
DB2CHGPWD_EEE DB2 UDB ESE on AIX Default=null
and Windows NT
Values: YES or NO
Specifies whether you allow other users to change passwords on AIX or Windows NT ESE systems. You must
ensure that the passwords for all partitions or nodes are maintained centrally using either a Windows NT domain
controller on Windows NT, or NIS on AIX. If not maintained centrally, passwords may not be consistent across all
partitions or nodes. This could result in a password being changed only at the database partition to which the user
connects to make the change. In order to modify this global registry variable, you must be at the root directory and
on the DAS instance.
This variable is required only if you use the old db2atld utility instead of the new LOAD utility.
DB2_FORCE_FCM_BP AIX Default=No
Values: Yes or No
This registry variable is applicable to DB2 UDB ESE for AIX with multiple logical partitions. When DB2START is
issued, DB2 allocates the FCM buffers either from the database global memory or from a separate shared memory
segment, if there is not enough global memory available. These buffers are used by all FCM daemons for that
instance on the same physical machine. The kind of memory allocated is largely dependent on the number of FCM
buffers to be created, as specified by the fcm_num_buffers database manager configuration parameter.
If the DB2_FORCE_FCM_BP variable is set to Yes, the FCM buffers are always created in a separate memory
segment so that communication between FCM daemons of different logical partitions on the same physical node
occurs through shared memory. Otherwise, FCM daemons on the same node communicate through UNIX Sockets.
Communicating through shared memory in is faster, but there is one fewer shared memory segment available for
other uses, particularly for database buffer pools. Enabling the DB2_FORCE_FCM_BP registry variable thus reduces
the maximum size of database buffer pools.
DB2_NUM_FAILOVER_NODES All Default: 2
For example, host A has two logical nodes: 1 and 2; and host B has two logical nodes: 3 and 4. Assume
DB2_NUM_FAILOVER_NODES is set to 2. During DB2START, both host A and host B will reserve enough memory
for FCM so that up to four logical nodes could be managed. Then if one host fails, the logical nodes for the failing
host could be restarted on the other host.
DB2_PARTITIONEDLOAD__DEFAULT All supported ESE Default: YES;
platforms
Range of values: YES/NO
The DB2_PARTITIONEDLOAD_DEFAULT registry variable lets users change the default behavior of the Load
utility in an ESE environment when no ESE-specific Load options are specified. The default value is YES, which
specifies that in an ESE environment if you do not specify ESE-specific Load options, loading is attempted on all
partitions on which the target table is defined.
When the value is NO, loading is attempted only on the partition to which the Load utility is currently connected.
DB2PORTRANGE Windows NT Values: nnnn:nnnn
This value is set to the TCP/IP port range used by FCM so that any additional partitions created on another
machine will also have the same port range.
In both ESE and NON-ESE environments, when EXTEND is specified, the optimizer searches for opportunities to
transform both ″NOT IN″ and ″NOT EXISTS″ subqueries into anti-joins.
DB2_CORRELATED_PREDICATES All Default=Yes
Values: Yes or No
The default for this variable is ″Yes″. When there are unique indexes on correlated columns in a join, and this
registry variable is ″Yes″, the optimizer attempts to detect and compensate for correlation of join predicates. When
this registry variable is ″Yes″, the optimizer uses the KEYCARD information of unique index statistics to detect
cases of correlation, and dynamically adjusts the combined selectivities of the correlated predicates, thus obtaining a
more accurate estimate of the join size and cost. Adjustment is also done for correlation of simple equality
predicates like WHERE C1=5 AND C2=10 if there is an index on C1 and C2. The index need not be unique but the
equality predicate columns must cover all the columns in the index.
DB2_HASH_JOIN All Default=YES
Values: YES or NO
Specifies hash join as a possible join method when compiling an access plan.
DB2_INLIST_TO_NLJN All Default=NO
Values: YES or NO
SELECT *
FROM EMPLOYEE
WHERE DEPTNO IN (’D11’, ’D21’, ’E21’)
This revision might provide better performance if there is an index on DEPTNO. The list of values would be
accessed first and joined to EMPLOYEE with a nested loop join using the index to apply the join predicate.
Sometimes the optimizer does not have accurate information to determine the best join method for the rewritten
version of the query. This can occur if the IN list contains parameter markers or host variables which prevent the
optimizer from using catalog statistics to determine the selectivity. This registry variable causes the optimizer to
favor nested loop joins to join the list of values, using the table that contributes the IN list as the inner table in the
join.
DB2_LIKE_VARCHAR All Default=Y,Y
Controls the use of sub-element statistics. These are statistics about the content of data in columns when the data
has a structure in the form of a series of sub-fields or sub-elements delimited by blanks. Collection of sub-element
statistics is optional and controlled by options in the RUNSTATS command or API.
This registry variable affects how the optimizer deals with a predicate of the form:
COLUMN LIKE ’%xxxxxx%’
where
v The term preceding the comma, or the only term to the right of the predicate, means the following but only if the
second term is specified as N or the column does not have positive sub-element statistics:
– S – The optimizer estimates the length of each element in a series of elements concatenated together to form a
column based on the length of the string enclosed in the % characters.
– Y – The default. Use a default value of 1.9 for the algorithm parameter. Use a variable-length sub-element
algorithm with the algorithm parameter.
– N – Use a fixed-length sub-element algorithm.
– num1 – Use the value of num1 as the algorithm parameter with the variable length sub-element algorithm.
v The term following the comma means the following, but only for columns that do have positive sub-element
statistics:
– N – Do not use sub-element statistics. The first term takes effect
– Y – The default. Use a variable-length sub-element algorithm that uses sub-element statistics together with the
1.9 default value for the algorithm parameter in the case of columns with positive sub-element statistics.
– num2 – Use a variable-length sub-element algorithm that uses sub-element statistics together with the value of
num2 as the algorithm parameter in the case of columns with positive sub-element statistics.
DB2_MINIMIZE_LISTPREFETCH All Default=NO
Values: YES or NO
This registry variable prevents the optimizer from considering list prefetch in such situations.
DB2_SELECTIVITY ALL Default=No
Values: Yes or No
This registry variable controls where the SELECTIVITY clause can be used in search conditions in SQL statements.
When this registry variable is set to ″Yes″, the SELECTIVITY clause can be specified for the following predicates:
v A basic predicate in which at least one expression contains host variables
v A LIKE predicate in which the MATCH expression, predicate expression, or escape expression contains host
variables
DB2_NEW_CORR_SQ_FF All Default=OFF
Values: ON or OFF
Affects the selectivity value computed by the SQL optimizer for certain subquery predicates when it is set to “ON”.
It can be used to improve the accuracy of the selectivity value of equality subquery predicates that use the MIN or
MAX aggregate function in the SELECT list of the subquery. For example:
SELECT * FROM T WHERE
[Link] = (SELECT MIN([Link])
FROM T WHERE ...)
DB2_PRED_FACTORIZE All Default=NO
Value: YES or NO
Specifies whether the optimizer searches for opportunities to extract additional predicates from disjuncts. In some
circumstances, the additional predicates can alter the estimated cardinality of the intermediate and final result sets.
With the following query:
SELECT [Link],
[Link]
FROM employee n1,
employee n2
WHERE
(([Link]=’SMITH’
AND [Link]=’JONES’)
OR ([Link]=’JONES’
AND [Link]=’SMITH’))
Note that the dynamic optimization reduction at optimization level 5 takes precedence over the behavior described
for optimization level of exactly 5 when DB2_REDUCED_OPTIMIZATION is set to YES as well as the behavior
described for the integer setting.
| DB2_SQLROUTINE_PREPOPTS All Default=empty string
| Values:
| v BLOCKING {UNAMBIG | ALL | NO}
| v DATETIME {DEF | USA | EUR | ISO |
| JIS | LOC}
| v DEGREE {1 | degree-of-parallelism | ANY}
| v DYNAMICRULES {BIND | RUN}
| v EXPLAIN {NO | YES | ALL}
| v EXPLSNAP {NO | YES | ALL}
| v FEDERATED {NO | YES}
| v INSERT {DEF | BUF}
| v ISOLATION {CS | RR | UR | RS | NC}
| v QUERYOPT optimization-level
| v VALIDATE {RUN | BIND}
| The DB2_SQLROUTINE_PREPOPTS registry variable can be used to customize the precompile and bind options for
| SQL procedures.
Related concepts:
v “Optimization class guidelines” on page 72
Appendix A. DB2 Registry and Environment Variables 505
v “Strategies for selecting optimal joins” on page 160
v “DB2 registry and environment variables” on page 489
Related reference:
v “Optimization classes” on page 73
Performance variables
Table 52. Performance Variables
Variable Name Operating System Values
Description
DB2AFFINITIES AIX 5 or higher, all Default=Not set
Linux except zSeries
(32–bit) Values: valid path to configuration file
Defines a resource policy which can be used to limit what operating system resources are used by DB2. For
example, on AIX or Linux, this registry variable can be used to limit the set of processors that DB2 uses.
On AIX NUMA enabled machines, a policy can be defined which specifies what resource sets DB2 will use. When
resource set binding is used, each individual DB2 process will be bound to a particular resource set. This can be
beneficial in some performance tuning scenarios.
The registry variable can be set to indicate the path to a configuration file which defines a policy for binding DB2
processes to operating system resources. The resource policy allows you to specify a set of operating system
resources to restrict DB2. Each DB2 process is bound to a single resource of the set. Resource assignment occurs in a
circular round robin fashion.
Note: Use of the RSET method requires CAP_NUMA_ATTACH capability and is not supported on Linux.
DB2_ALLOCATION_SIZE All Default=8 MB
Range: 32 KB–256 MB
| Specifies the size of memory allocations for buffer pools.
| The potential advantage of setting a higher value for this registry variable is that it will require fewer allocations to
| reach a desired amount of memory that is allocated to a buffer pool.
| The potential cost of setting a higher value for this registry variable is that memory can be wasted if the buffer pool
| is altered by a non-multiple of the allocation size. For example, if the value for DB2_ALLOCATION_SIZE is 8 MB
| and a buffer pool is reduced by 4 MB, this 4 MB will be wasted because an entire 8 MB segment cannot be freed.
Values: ON , OFF
Set this variable to ON to enable performance-related changes in the access plan manager (APM) that affect the
behavior of the SQL cache (package cache). These settings are not usually recommended for production systems.
They introduce some limitations, such as the possibility of out-of-package cache errors or increased memory use or
both.
Setting DB2_APM_PERFORMANCE to ON also enables the ’No Package Lock’ mode. This mode allows the Global
SQL Cache to operate without the use of package locks, which are internal system locks that protect cached package
entries from being removed. The ’No Package Lock’ mode might result in somewhat improved performance, but
certain database operations are not allowed. These prohibited operations might include: operations that invalidate
packages, operations that inoperate packages, and PRECOMPILE, BIND, and REBIND.
DB2ASSUMEUPDATE All Default=OFF
Values: ON or OFF
Specifies whether prefetch should be used during crash recovery. If DB2_AVOID_PREFETCH=ON, prefetch is not
used.
DB2_AWE Windows 2000 Default=null
Values: YES or NO
Parameters are specified in an ASCII file, one parameter on each line, in the form parameter=value. For example, a
file named [Link] might contain the following lines:
NO_NT_SCATTER = 1
NUMPREFETCHQUEUES = 2
Assuming that [Link] is stored in F:\vars\, to set these variables you execute the following command:
db2set DB2BPVARS=F:\vars\[Link]
Scatter-read parameters
The scatter-read parameters are recommended for systems with a large amount of sequential prefetching against the
respective type of containers and for which you have already set DB2NTNOCACHE to ON. These parameters,
available only on Windows platforms, are NT_SCATTER_DMSFILE, NT_SCATTER_DMSDEVICE, and
NT_SCATTER_SMS. Specify the NO_NT_SCATTER parameter to explicitly disallow scatter read for any container.
Specific parameters are used to turn scatter read on for all containers of the indicated type. For each of these
parameters, the default is zero (or OFF); and the possible values include: zero (or OFF) and 1 (or ON).
Note: You can turn on scatter read only if DB2NTNOCACHE is set to ON to turn Windows file caching off. If
DB2NTNOCACHE is set to OFF or not set, a warning message is written to the administration notification log if
you attempt to turn on scatter read for any container, and scatter read remains disabled.
Prefetch-adjustment parameters
If you think the default values are too small for your environment, first increase the values only slightly. For
example, you might set NUMPREFETCHQUEUES=4 and PREFETCHQUEUESIZE=200. Make changes to these
parameters in a controlled manner so that you can monitor and evaluate the effects of the change.
For NUMPREFETCHQUEUES, the default is 1, and the range of values is 1 to NUM_IOSERVERS. If you set
NUMPREFETCHQUEUES to less than 1, it is adjusted to 1. If you set it greater than NUM_IOSERVERS, it is
adjusted to NUM_IOSERVERS.
For PREFETCHQUEUESIZE, the default value is max(100,2*NUM_IOSERVERS). The range of values is 1 to 32767. If
you set PREFETCHQUEUESIZE to less than 1, it is adjusted to the default. If set greater than 32767, it is adjusted to
32767.
Values: ON or OFF
Specifies whether or not pointer checking for input is required.
| DB2CHKSQLDA All Default=OFF,
Values: ON or OFF
| Specifies whether or not SQLDA checking for input is required.
| DB2_ENABLE_BUFPD All Default=YES
Values: ON or OFF
Specifies whether or not DB2 uses intermediate buffering to improve query performance. The buffering may not
improve query performance in all environments. Testing should be done to determine individual query performance
improvements.
DB2_EVALUNCOMMITTED All Default=OFF
With this variable enabled, predicate evaluation may occur on uncommitted data.
It is applicable only to statements using either Cursor Stability or Read Stability isolation levels. For index scans, the
index must be a type-2 index.
Furthermore, deleted rows are skipped unconditionally on table scan access while deleted keys are not skipped for
type-2 index scans unless the registry variable DB2_SKIPDELETED is also set.
The activation of this the DB2_EVALUNCOMMITTED registry variable is effective on db2start while the decision as
to whether deferred locking is applicable, is made at statement compile or bind time.
DB2_EXTENDED_OPTIMIZATION All Default=OFF
Values: ON or OFF
Specifies whether or not the query optimizer uses optimization extensions to improve query performance. The
extensions may not improve query performance in all environments. Testing should be done to determine
individual query performance improvements.
| DB2_KEEPTABLELOCK All Default=OFF
| On 64-bit DC2 for AIX, enabling this variable will reduce the size of the shared memory segment backing database
| memory to the minimum requirement (the default is to create a 64GB segment - see the database_memory
| configuration parameter for more details). This is to avoid pinning more shared memory in RAM than is likely to
| be used.
| With this variable set, the ability to dynamically increase the overall database shared memory configuration, for
| example, to increase the size of buffer pools, will be limited.
| On Linux, there is an additional requirement for the availablility of the [Link] library. This library must be
| installed for this option to work. If this option is turned on, and the library is not on the system, DB2 will disable
| the large Kernel pages and continue to function as it would previously.
| On Linux, to verify that Large Kernel Pages are available, issue the following command:
| cat /proc/meminfo
| If it is available, the following three lines should appear (with different numbers depending on the amount of
| memory configured on your machine)
| HugePages_Total: 200
| HugePages_Free: 200
| Hugepagesize: 16384 kB
| If you do not see these lines, or if the HugePages_Total is 0, configuration of the operating system or kernel is
| required.
DB2MAXFSCRSEARCH All Default=5
| For best results, the recommended value for this variable is the maximum number of tables expected to be accessed
| by any connections. If no user-defined value is specified, the default value is as follows: If the locklist size is greater
| than or equal to SQLP_THRESHOLD_VAL_OF_LRG_LOCKLIST_SZ_FOR_MAX_NON_LOCKS (currently 8000), the
| default value will be SQLP_DEFAULT_MAX_NON_TABLE_LOCKS_LARGE (currently 150). Otherwise the default
| value will be SQLP_DEFAULT_MAX_NON_TABLE_LOCKS_SMALL (currently 0).
DB2MEMDISCLAIM AIX Default=YES
Values: YES or NO
A DB2MEMDISCLAIM setting of YES results in smaller paging space requirements, and possibly less disk activity
from paging. A DB2MEMDISCLAIM setting of NO will result in larger paging space requirements, and possibly
more disk activity from paging. In some situations, such as if paging space is plentiful and real memory is so
plentiful that paging never occurs, a setting of NO provides a minor performance improvement.
DB2MEMMAXFREE All Default= 8 388 608 bytes
Values: ON or OFF
Used in conjunction with DB2_MMAP_WRITE to allow DB2 to use mmap as an alternate method of I/O. In most
environments, mmap should be used to avoid operating system locks when multiple processes are writing to
different sections of the same file.
When these variables are set to ON, data that is read to and written from the DB2 buffer pools bypasses the AIX
memory cache. If you have a relatively small DB2 buffer pool, and you cannot or choose not to increase the size of
this buffer pool, you should consider taking advantage of AIX memory caching by setting DB2_MMAP_READ and
DB2_MMAP_WRITE to OFF.
DB2_MMAP_WRITE AIX Default=ON
Values: ON or OFF
Used in conjunction with DB2_MMAP_READ to allow DB2 to use mmap as an alternate method of I/O. In most
environments, mmap should be used to avoid operating system locks when multiple processes are writing to
different sections of the same file.
When these variables are set to ON, data that is read to and written from the DB2 buffer pools bypasses the AIX
memory cache. If you have a relatively small DB2 buffer pool, and you cannot or choose not to increase the size of
this buffer pool, you should consider taking advantage of AIX memory caching by setting DB2_MMAP_READ and
DB2_MMAP_WRITE to OFF.
| DB2_NO_FORK_CHECK UNIX Default=OFF
Values: YES
| Databases created by the CREATE DATABASE command or the equivalent API have multipage file allocation
| (MPFA) enabled. Once MPFA is enabled for a database, it cannot be disabled. To create a database with MPFA
| disabled, set this registry variable to YES and restart the instance before creating the database. When this registry
| variable is set, any created database will have MPFA disabled.
| To enable MPFA for a database that has MPFA disabled, use the db2empfa command.
DB2NTMEMSIZE Windows NT Default=(varies by memory segment)
Value: ON or OFF
Specifies whether DB2 opens database files with a NOCACHE option. If DB2NTNOCACHE=ON, file system
caching is eliminated. If DB2NTNOCACHE=OFF, the operating system caches DB2 files. This applies to all data
except for files that contain long fields or LOBs. Eliminating system caching allows more memory to be available to
the database so that the buffer pool or sortheap can be increased.
In Windows NT, files are cached when they are opened, which is the default behavior. 1 MB is reserved from a
system pool for every 1 GB in the file. Use this registry variable to override the undocumented 192 MB limit for the
cache. When the cache limit is reached, an out-of-resource error is given.
DB2NTPRICLASS Windows NT Default=null
This variable is used in conjunction with individual thread priorities (set using DB2PRIORITIES) to determine the
absolute priority of DB2 threads relative to other threads in the system.
Note: Care should be taken when using this variable. Misuse could adversely affect overall system performance.
For more information, please refer to the SetPriorityClass() API in the Win32 documentation.
DB2NTWORKSET Windows NT Default=1,1
Used to modify the minimum and maximum working-set size available to DB2. By default, when Windows NT is
not in a paging situation, the working set of a process can grow as large as needed. However, when paging occurs,
the maximum working set that a process can have is approximately 1 MB. DB2NTWORKSET allows you to override
this default behavior.
Specify DB2NTWORKSET for DB2 using the syntax DB2NTWORKSET=min,max, where min and max are expressed
in megabytes.
DB2_OVERRIDE_BPF All Default=not set
You can also use <entry>[;<entry>...] where <entry>=<buffer pool ID>,<number of pages> to temporarily change
the size of all or a subset of the buffer pools so that they can start up.
DB2_PINNED_BP AIX, HP-UX Default=NO
Values: YES or NO
This variable is used to specify the database global memory (including buffer pools) associated with the database in
the main memory on some AIX operating systems. Keeping this database global memory in the system main
memory allows database performance to be more consistent.
For example, if the buffer pool is swapped out of the system main memory, database performance deteriorates. The
reduction of disk I/O by having the buffer pools in system memory improves database performance. If other
applications require more of the main memory, allow the database global memory to be swapped out of main
memory, depending on the system main memory requirements.
| On 64-bit DB2 for AIX, enabling this variable will reduce the size of the shared memory segment backing database
| memory to the minimum requirement (the default is to create a 64GB segment - see the database_memory
| configuration parameter for more details). This is to avoid pinning more shared memory in RAM than is likely to
| be used.
| With this variable set, the ability to dynamically increase the overall database shared memory configuration, for
| example, to increase the size of buffer pools, will be limited.
For HP-UX in a 64-bit environment, in addition to modifying this registry variable, the DB2 instance group must be
given the MLOCK privilege. To do this, a user with root access rights performs the following actions:
1. Adds the DB2 instance group to the /etc/privgroup file. For example, if the DB2 instance group belongs to
db2iadm1 group then the following line must be added to the /etc/privgroup file:
db2iadm1 MLOCK
2. Issues the following command:
setprivgrp -f /etc/privgroup
This kernel patch is currently in UnitedLinux 1.0 SP2 or higher for IA-32 and will be in all upcoming Linux 2.6
kernels.
DB2_SKIPDELETED All Default=OFF
This registry variable does not impact the behavior of cursors on the DB2 catalog tables.
If this variable is set to -1, then the file is not be truncated at all and the file will be allowed to grow indefinitely,
restricted only by system resources.
If this variable is set to 0, then no special threshold handling is done. Instead, once a temporary table is no longer
needed, that file is truncated to 0.
DB2_SORT_AFTER_TQ All Default=NO
Values: YES or NO
Specifies how the optimizer works with directed table queues in a partitioned database when the receiving end
requires the data to be sorted and the number of receiving nodes is equal to the number of sending nodes.
When DB2_SORT_AFTER_TQ= NO, the optimizer tends to sort at the sending end and merge the rows at the
receiving end.
When DB2_SORT_AFTER_TQ= YES, the optimizer tends to transmit the rows unsorted, not merge at the receiving
end, and sort the rows at the receiving end after receiving all the rows.
| DB2_SELUDI_COMM_BUFFER All Default=OFF
Values=ON, OFF
| Applies to the processing of blocking cursors over SELECT from UPDATE/INSERT/DELETE (UDI) queries. When
| enabled, this registry variable prevents the result of a query from being stored in a temporary table. Instead, during
| the OPEN processing of a blocking cursor for a SELECT from UDI query, DB2 attempts to buffer the entire result of
| the query directly into the communications buffer memory area.
| Note: If the communications buffer space is not large enough to hold the entire result of query, SQLCODE -906 is
| issued and the transaction is rolled back. See the aslheapsz and rqrioblk database manager configuration parameters
| for information on adjusting the size of the communication buffer memory area for local and remote applications
| respectively.
| This registry variable is not supported in partitioned database environments or when intra-partition parallelism is
| enabled.
DB2_TRUSTED_BINDIN All Default=OFF
When this variable is enabled, there is no conversion from the external SQLDA format to an internal DB2 format
during the binding of SQL statements contained within an embedded unfenced stored procedure. This will speed
up the processing of the embedded SQL statements.
The following datatypes are not supported in embedded unfenced stored procedures when this variable is enabled:
v SQL_TYP_DATE
v SQL_TYP_TIME
v SQL_TYP_STAMP
v SQL_TYP_DATALINK
v SQL_TYP_CGSTR
v SQL_TYP_BLOB
v SQL_TYP_CLOB
v SQL_TYP_DBCLOB
v SQL_TYP_CSTR
v SQL_TYP_LSTR
v SQL_TYP_BLOB_LOCATOR
v SQL_TYP_CLOB_LOCATOR
v SQL_TYP_DCLOB_LOCATOR
v SQL_TYP_BLOB_FILE
v SQL_TYP_CLOB_FILE
v SQL_TYP_DCLOB_FILE
v SQL_TYP_BLOB_FILE_OBSOLETE
v SQL_TYP_CLOB_FILE_OBSOLETE
v SQL_TYP_DCLOB_FILE_OBSOLETE
If these datatypes are encountered, an SQLCODE −804, SQLSTATE 07002 will be returned.
Note: The data type and length of the input host variable has to match exactly the internal data type and length of
the corresponding element. For host variables, this requirement will always be met. However, for parameter
markers, care must be taken to ensure that matching data types are used. The CHECK option can be used to ensure
that the data types and lengths match for all input host variables, but this option negates most of the performance
improvements.
DB2_USE_ALTERNATE_PAGE_CLEANING All Default=not set
Related concepts:
v “DB2 registry and environment variables” on page 489
If this variable is set to XBSA, you must set the DLFM_BACKUP_TARGET_LIBRARY variable.
If you change the setting of this registry variable from one target to another, the archived files are not moved. Only
new backups are placed in the new location. Previously archived files are not moved.
DLFM_BACKUP_TARGET_LIBRARY AIX, Windows NT, Default: null
Windows 2000,
Solaris Operating Values: any valid path to the DLL or shared
Environment library name
Specifies the fully qualified path to the XBSA-compliant archive server DLL or shared library. This library is loaded
using the libdfmxbsa.a library.
This variable must be set if the DLFM_BACKUP_TARGET is set to XBSA. It does not apply if the
DLFM_BACKUP_TARGET variable is set to another value.
DLFM_GC_MODE AIX, Windows NT, Default: PASSIVE
Windows 2000,
Solaris Operating Values: SLEEP, PASSIVE, or ACTIVE
Environment
Specifies the control of garbage file collection on the Data Links server. When set to SLEEP, no garbage collection
occurs. When set to PASSIVE, garbage collection runs only if no other transactions are running. When set to
ACTIVE, garbage collection runs even if other transactions are running.
DLFM_INSTALL_PATH AIX, Windows NT, Default
Windows 2000,
Solaris Operating On AIX and the Solaris Operating
Environment Environment: /home/<instance>/sqllib/bin
where <instance> is the Data Links Manager
instance ID
When this variable is set to YES, the DLFM_ASNCOPYD_PORT variable must be set.
DLFM_ASNCOPYD_PORT AIX, Windows NT, Default: null
Windows 2000,
Solaris Operating Values: any valid port number
Environment
Specifies the TCP/IP port number on which the Data Links Manager Replication Daemon (DLFM_ASNCOPYD) will
listen for file replication requests.
This variable must be specified when the DLFM_START_ASNCOPYD variable is set to YES.
DLFM_NUM_ARCHIVE_SUBSYSTEMS AIX, Windows NT, Default: 2
Windows 2000,
Solaris Operating Values: any number greater than or equal to 1
Environment
Specifies the number of DLFM Copy Daemon processes to run under a given DLFM server. The larger the number
of copy processes, the greater the throughput on backing up linked files. However, this value should correspond to
the number of I/O channels available for copying linked files to the designated archive area. If the value is too
large, the amount of system resources consumed can reduce the benefits of I/O parallelism.
DLFM_AUTOSTART AIX, Solaris Default: NO
Operating
Environment Values: YES, NO
Specifies whether the DLFM server is automatically started whenever the operating system reboots. This variable is
checked by the dlfsmount script, invoked from the /etc/inittab file during boot processing.
Related concepts:
v “DB2 registry and environment variables” on page 489
| Note that this registry variable will be deprecated in version 10, and the commit-on-exit behavior will no longer be
| supported. Users should determine whether any of their applications developed prior to version 8 continue to
| depend on this functionality, and add the appropriate explicit COMMIT statements to the application as required. If
| the registry variable is turned on, care should be taken not to implement new applications which fail to explicitly
| COMMIT before exit.
| Most users should leave this registry variable at the default setting.
DB2DEFPREP All Default=NO
| This variable is ignored unless the database manager parameter FEDERATED is set to YES.
| DB2_DJ_INI All Default:
v UNIX:
db2_instance_directory/cfg/[Link]
v Windows:
db2_install_directory\cfg\[Link]
| This variable is ignored unless the database manager parameter FEDERATED is set to YES.
DB2DMNBCKCTLR Windows NT Default=null
| Values: ON or OFF
| Prevents unauthorized access to DB2 by locking DB2 system files. To avoid potential problems, this registry varible
| should not be turned off.
Values: YES or NO
Specifies whether or not the Lightweight Directory Access Protocol (LDAP) is used. LDAP is an access method to
directory services.
DB2_FALLBACK Windows NT Default=OFF
Values: ON or OFF
This variable allows you to force all database connections off during the fallback processing. It is used in
conjunction with the failover support in the Windows NT environment with Microsoft Cluster Server (MSCS). If
DB2_FALLBACK is not set or is set to OFF, and a database connection exists during the fall back, the DB2 resource
cannot be brought offline. This will mean the fallback processing will fail.
DB2_FMP_COMM_HEAPSZ Windows, all UNIX 20mb or enough space to run 10 fenced
except AIX routines (whichever is larger)
This variable specifies, in 4 KB pages, the size of the pool used for fenced routine invocations, such as stored
procedure or user-defined function calls. The space used by each fenced routine is twice the value of the aslheapsz
configuration parameter.
If you are running a large number of fenced routines on your system, you may need to increase the value of this
variable. If you are running a very small number of fenced routines, you can reduce it.
| Setting this value to 0 means that no set is created, and as a result no fenced routines can be invoked. It also means
| that the health monitor and the automatic database maintenance functionality (such as automatic backups, statistics
| collection, and REORG) will be disabled since this functionality relies on the fenced routine infrastructure.
DB2_GRP_LOOKUP Windows NT Default=null
| If HADR synchronization mode (the HADR_SYNCMODE database configuration parameter) is set to ASYNC,
| during peer state, a slow standby may cause the send operation on the primary to stall and therefore block
| transaction processing on the primary. A larger than default log-receiving buffer can be configured on a standby
| database to allow it to hold more unprocessed log data. This may allow for brief periods where the primary
| generates log data faster than the standby can consume it, without blocking transaction processing at the primary.
DB2LDAP_BASEDN All Default=null
Values: YES or NO
To ensure that you have the latest entries in the cache, do the following:
REFRESH LDAP DB DIR
REFRESH LDAP NODE DIR
These commands update and remove incorrect entries from the database directory and the node directory.
DB2LDAP_CLIENT_PROVIDER Windows Default=null (Microsoft, if available, is used;
otherwise IBM is used.)
Values: YES, NO
Specifies whether DB2 caches its internal LDAP connection handles. When this variable is set to NO, DB2 will not
cache its LDAP connection handles to the directory server. This will likely result in a negative performance impact,
but it might be desirable to set DB2LDAP_KEEP_CONNECTION to NO if the number of simultaneously active
LDAP client connections to the directory server needs to be minimized.
The DB2LDAP_KEEP_CONNECTION registry variable is only implemented as a global level profile registry
variable in LDAP, so you must set it by specifying the -gl option with the db2set command as follows:
db2set -gl DB2LDAP_KEEP_CONNECTION=NO
DB2LDAP_SEARCH_SCOPE All Default= DOMAIN
Values: STATEMENT
Specifies whether lock timeouts cause the entire transaction to be rolled back, or only the current statement. If
DB2LOCK_TO_RB is set to STATEMENT, locked timeouts cause only the current statement to be rolled back. Any other
setting results in transaction rollback.
DB2_NEWLOGPATH2 UNIX Default=0
Values: 0 or 1
This parameter allows you to specify whether a secondary path should be used to implement dual logging. The
secondary path name is generated by appending a “2” to the current value of the logpath database configuration
parameter.
DB2NOEXITLIST All Default=OFF
Values: ON or OFF
If defined, this variable indicates to DB2 not to install an exit list handler in applications and not to perform a
COMMIT. Normally, DB2 installs a process exit list handler in applications and the exit list handler performs a
COMMIT operation if the application ends normally.
For applications that dynamically load the DB2 library and unload it before the application terminates, the
invocation of the exit list handler fails because the handler routine is no longer loaded in the application. If your
application operates in this way, you should set the DB2NOEXITLIST variable and ensure your application
explicitly invokes all required COMMITs.
DB2OLDEVMON All Values: event monitor names separated by a
commaevmon1, evmon2, ...
Specifies the names of event monitors that write data in the pre-Version 6 format. In DB2 Version 6, the
self-describing data stream became the standard form of system monitor output to files and pipes. Pre-Version 6,
system monitor data was returned in fixed data structures.
DB2REMOTEPREG Windows NT Default=null
This name is displayed in the system level of the Control Center’s object tree to aid administrators in the
identification of server systems that can be administered from the Control Center.
When using the ’Search the Network’ function of the Client Configuration Assistant, DB2 discovery returns this
name and it is displayed at the system level in the resulting object tree. This name aids users in identifying the
system that contains the database they wish to access. A value for DB2SYSTEM is set at installation time as follows:
v On Windows NT the setup program sets it equal to the computer name specified for the Windows system.
v On UNIX systems, it is set equal to the UNIX system’s TCP/IP hostname.
DB2_VENDOR_INI AIX, HP-UX, the Default=null
Solaris Operating
Environment, and Values: Any valid path and file.
Windows
Points to a file containing all vendor-specific environment settings. The value is read when the database manager
starts.
DB2_XBSA_LIBRARY AIX, HP-UX, the Default=null
Solaris Operating
Environment,, and Values: Any valid path and file.
Windows
Points to the vendor-supplied XBSA library. On AIX, the setting must include the shared object if it is not named
shr.o. HP-UX, the Solaris Operating Environment, and Windows NT do not require the shared object name. For
example, to use Legato’s NetWorker Business Suite Module for DB2, the registry variable must be set as follows:
db2set DB2_XSBA_LIBRARY="/usr/lib/libxdb2.a(bsashr10.o)"
The XBSA interface can be invoked through the BACKUP DATABASE or the RESTORE DATABASE commands. For
example:
db2 backup db sample use XBSA
db2 restore db sample use XBSA
Related concepts:
v “DB2 registry and environment variables” on page 489
The population of the Explain tables by the Explain facility will not activate
triggers or referential or check constraints. For example, if an insert trigger were
defined on the EXPLAIN_INSTANCE table, and an eligible statement were
explained, the trigger would not be activated.
Related reference:
v “EXPLAIN_ARGUMENT table” on page 526
v “EXPLAIN_OBJECT table” on page 532
v “EXPLAIN_OPERATOR table” on page 535
v “EXPLAIN_PREDICATE table” on page 537
v “EXPLAIN_STREAM table” on page 541
v “ADVISE_INDEX table” on page 543
v “ADVISE_WORKLOAD table” on page 550
v “EXPLAIN_INSTANCE table” on page 530
v “EXPLAIN_STATEMENT table” on page 539
v “ADVISE_INSTANCE table” on page 546
v “ADVISE_MQT table” on page 547
v “ADVISE_PARTITION table” on page 548
v “ADVISE_TABLE table” on page 549
EXPLAIN_ARGUMENT table
The EXPLAIN_ARGUMENT table represents the unique characteristics for each
individual operator, if there are any.
Table 55. EXPLAIN_ARGUMENT Table. PK means that the column is part of a primary key; FK means that the
column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
EXPLAIN_REQUESTER VARCHAR(128) No FK Authorization ID of initiator of this Explain request.
EXPLAIN_TIME TIMESTAMP No FK Time of initiation for Explain request.
SOURCE_NAME VARCHAR(128) No FK Name of the package running when the dynamic
statement was explained or name of the source file
when static SQL was explained.
SOURCE_SCHEMA VARCHAR(128) No FK Schema, or qualifier, of source of Explain request.
SOURCE_VERSION VARCHAR(64) No FK Version of the source of the Explain request.
EXPLAIN_LEVEL CHAR(1) No FK Level of Explain information for which this row is
relevant.
STMTNO INTEGER No FK Statement number within package to which this
explain information is related.
SECTNO INTEGER No FK Section number within package to which this explain
information is related.
OPERATOR_ID INTEGER No No Unique ID for this operator within this query.
ARGUMENT_TYPE CHAR(8) No No The type of argument for this operator.
ARGUMENT_VALUE VARCHAR(1024) Yes No The value of the argument for this operator. NULL if
the value is in LONG_ARGUMENT_VALUE.
LONG_ARGUMENT_VALUE CLOB(2M) Yes No The value of the argument for this operator, when the
text will not fit in ARGUMENT_VALUE. NULL if the
value is in ARGUMENT_VALUE.
(A) Ascending
(D) Descending
EXPLAIN_INSTANCE table
The EXPLAIN_INSTANCE table is the main control table for all Explain
information. Each row of data in the Explain tables is explicitly linked to one
unique row in this table. The EXPLAIN_INSTANCE table gives basic information
about the source of the SQL statements being explained as well as information
about the environment in which the explanation took place.
Table 57. EXPLAIN_INSTANCE Table. PK means that the column is part of a primary key; FK means that the column
is part of a foreign key.
Column Name Data Type Nullable? Key? Description
EXPLAIN_REQUESTER VARCHAR(128) No PK Authorization ID of initiator of this Explain request.
EXPLAIN_TIME TIMESTAMP No PK Time of initiation for Explain request.
SOURCE_NAME VARCHAR(128) No PK Name of the package running when the dynamic
statement was explained or name of the source file
when the static SQL was explained.
SOURCE_SCHEMA VARCHAR(128) No PK Schema, or qualifier, of source of Explain request.
SOURCE_VERSION VARCHAR(64) No PK Version of the source of the Explain request.
EXPLAIN_OPTION CHAR(1) No No Indicates what Explain Information was requested for
this request.
Table 57. EXPLAIN_INSTANCE Table (continued). PK means that the column is part of a primary key; FK means that
the column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
ISOLATION CHAR(2) No No Indicates what type of isolation was used when
compiling the SQL statements. For more information,
see the ISOLATION column in [Link].
EXPLAIN_OBJECT table
The EXPLAIN_OBJECT table identifies those data objects required by the access
plan generated to satisfy the SQL statement.
Table 58. EXPLAIN_OBJECT Table. PK means that the column is part of a primary key; FK means that the column is
part of a foreign key.
Column Name Data Type Nullable? Key? Description
EXPLAIN_REQUESTER VARCHAR(128) No FK Authorization ID of initiator of this Explain request.
EXPLAIN_TIME TIMESTAMP No FK Time of initiation for Explain request.
SOURCE_NAME VARCHAR(128) No FK Name of the package running when the dynamic
statement was explained or name of the source file when
the static SQL was explained.
SOURCE_SCHEMA VARCHAR(128) No FK Schema, or qualifier, of source of Explain request.
SOURCE_VERSION VARCHAR(64) No FK Version of the source of the Explain request.
EXPLAIN_LEVEL CHAR(1) No FK Level of Explain information for which this row is
relevant.
STMTNO INTEGER No FK Statement number within package to which this explain
information is related.
SECTNO INTEGER No FK Section number within package to which this explain
information is related.
OBJECT_SCHEMA VARCHAR(128) No No Schema to which this object belongs.
OBJECT_NAME VARCHAR(128) No No Name of the object.
OBJECT_TYPE CHAR(2) No No Descriptive label for the type of object.
CREATE_TIME TIMESTAMP Yes No Time of Object’s creation; null if a table function.
STATISTICS_TIME TIMESTAMP Yes No Last time of update to statistics for this object; null if
statistics do not exist for this object.
COLUMN_COUNT SMALLINT No No Number of columns in this object.
ROW_COUNT INTEGER No No Estimated number of rows in this object.
WIDTH INTEGER No No The average width of the object in bytes. Set to -1 for an
index.
PAGES INTEGER No No Estimated number of pages that the object occupies in
the buffer pool. Set to -1 for a table function.
DISTINCT CHAR(1) No No Indicates whether the rows in the object are distinct (that
is, whether there are duplicates).
Table 58. EXPLAIN_OBJECT Table (continued). PK means that the column is part of a primary key; FK means that
the column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
CLUSTER DOUBLE No No Degree of data clustering with the index. If >= 1, this is
the CLUSTERRATIO. If >= 0 and < 1, this is the
CLUSTERFACTOR. Set to -1 for a table, table function,
or if this statistic is not available.
NLEAF INTEGER No No Number of leaf pages this index object’s values occupy.
Set to -1 for a table, table function, or if this statistic is
not available.
NLEVELS INTEGER No No Number of index levels in this index object’s tree. Set to
-1 for a table, table function, or if this statistic is not
available.
FULLKEYCARD BIGINT No No Number of distinct full key values contained in this
index object. Set to -1 for a table, table function, or if this
statistic is not available.
OVERFLOW INTEGER No No Total number of overflow records in the table. Set to -1
for an index, table function, or if this statistic is not
available.
FIRSTKEYCARD BIGINT No No Number of distinct first key values. Set to −1 for a table,
table function, or if this statistic is not available.
FIRST2KEYCARD BIGINT No No Number of distinct first key values using the first {2,3,4}
columns of the index. Set to −1 for a table, table
FIRST3KEYCARD BIGINT No No
function, or if this statistic is not available.
FIRST4KEYCARD BIGINT No No
SEQUENTIAL_PAGES INTEGER No No Number of leaf pages located on disk in index key order
with few or no large gaps between them. Set to −1 for a
table, table function, or if this statistic is not available.
DENSITY INTEGER No No Ratio of SEQUENTIAL_PAGES to number of pages in
the range of pages occupied by the index, expressed as a
percentage (integer between 0 and 100). Set to −1 for a
table, table function, or if this statistic is not available.
STATS_SRC CHAR(1) No No Indicates the source for the statistics. Set to 1 if from
single node.
AVERAGE_SEQUENCE_ DOUBLE No No Gap between sequences.
GAP
AVERAGE_SEQUENCE_ DOUBLE No No Gap between sequences when fetching using the index.
FETCH_GAP
AVERAGE_SEQUENCE_ DOUBLE No No Average number of index pages accessible in sequence.
PAGES
AVERAGE_SEQUENCE_ DOUBLE No No Average number of table pages accessible in sequence
FETCH_PAGES when fetching using the index.
AVERAGE_RANDOM_ DOUBLE No No Average number of random index pages between
PAGES sequential page accesses.
AVERAGE_RANDOM_ DOUBLE No No Average number of random table pages between
FETCH_PAGES sequential page accesses when fetching using the index.
NUMRIDS BIGINT No No Total number of row identifiers in the index.
NUMRIDS_DELETED BIGINT No No Total number of psuedo-deleted row identifiers in the
index.
NUM_EMPTY_LEAFS BIGINT No No Total number of empty leaf pages in the index.
ACTIVE_BLOCKS BIGINT No No Total number of active multidimensional clustering
(MDC) blocks in the table.
EXPLAIN_OPERATOR table
The EXPLAIN_OPERATOR table contains all the operators needed to satisfy the
SQL statement by the SQL compiler.
Table 60. EXPLAIN_OPERATOR Table. PK means that the column is part of a primary key; FK means that the
column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
EXPLAIN_REQUESTER VARCHAR(128) No FK Authorization ID of initiator of this Explain request.
EXPLAIN_TIME TIMESTAMP No FK Time of initiation for Explain request.
SOURCE_NAME VARCHAR(128) No FK Name of the package running when the dynamic
statement was explained or name of the source file when
the static SQL was explained.
SOURCE_SCHEMA VARCHAR(128) No FK Schema, or qualifier, of source of Explain request.
SOURCE_VERSION VARCHAR(64) No FK Version of the source of the Explain request.
EXPLAIN_LEVEL CHAR(1) No FK Level of Explain information for which this row is
relevant.
STMTNO INTEGER No FK Statement number within package to which this explain
information is related.
SECTNO INTEGER No FK Section number within package to which this explain
information is related.
OPERATOR_ID INTEGER No No Unique ID for this operator within this query.
OPERATOR_TYPE CHAR(6) No No Descriptive label for the type of operator.
TOTAL_COST DOUBLE No No Estimated cumulative total cost (in timerons) of
executing the chosen access plan up to and including
this operator.
IO_COST DOUBLE No No Estimated cumulative I/O cost (in data page I/Os) of
executing the chosen access plan up to and including
this operator.
CPU_COST DOUBLE No No Estimated cumulative CPU cost (in instructions) of
executing the chosen access plan up to and including
this operator.
FIRST_ROW_COST DOUBLE No No Estimated cumulative cost (in timerons) of fetching the
first row for the access plan up to and including this
operator. This value includes any initial overhead
required.
RE_TOTAL_COST DOUBLE No No Estimated cumulative cost (in timerons) of fetching the
next row for the chosen access plan up to and including
this operator.
RE_IO_COST DOUBLE No No Estimated cumulative I/O cost (in data page I/Os) of
fetching the next row for the chosen access plan up to
and including this operator.
RE_CPU_COST DOUBLE No No Estimated cumulative CPU cost (in instructions) of
fetching the next row for the chosen access plan up to
and including this operator.
COMM_COST DOUBLE No No Estimated cumulative communication cost (in TCP/IP
frames) of executing the chosen access plan up to and
including this operator.
FIRST_COMM_COST DOUBLE No No Estimated cumulative communications cost (in TCP/IP
frames) of fetching the first row for the chosen access
plan up to and including this operator. This value
includes any initial overhead required.
BUFFERS DOUBLE No No Estimated buffer requirements for this operator and its
inputs.
REMOTE_TOTAL_COST DOUBLE No No Estimated cumulative total cost (in timerons) of
performing operation(s) on remote database(s).
Table 60. EXPLAIN_OPERATOR Table (continued). PK means that the column is part of a primary key; FK means
that the column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
REMOTE_COMM_COST DOUBLE No No Estimated cumulative communication cost of executing
the chosen remote access plan up to and including this
operator.
EXPLAIN_PREDICATE table
The EXPLAIN_PREDICATE table identifies which predicates are applied by a
specific operator.
Table 62. EXPLAIN_PREDICATE Table. PK means that the column is part of a primary key; FK means that the
column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
EXPLAIN_REQUESTER VARCHAR(128) No FK Authorization ID of initiator of this Explain request.
EXPLAIN_TIME TIMESTAMP No FK Time of initiation for Explain request.
SOURCE_NAME VARCHAR(128) No FK Name of the package running when the dynamic
statement was explained or name of the source file when
the static SQL was explained.
SOURCE_SCHEMA VARCHAR(128) No FK Schema, or qualifier, of source of Explain request.
SOURCE_VERSION VARCHAR(64) No FK Version of the source of the Explain request.
EXPLAIN_LEVEL CHAR(1) No FK Level of Explain information for which this row is
relevant.
STMTNO INTEGER No FK Statement number within package to which this explain
information is related.
SECTNO INTEGER No FK Section number within package to which this explain
information is related.
OPERATOR_ID INTEGER No No Unique ID for this operator within this query.
PREDICATE_ID INTEGER No No Unique ID for this predicate for the specified operator.
HOW_APPLIED CHAR(5) No No How predicate is being used by the specified operator.
WHEN_EVALUATED CHAR(3) No No Indicates when the subquery used in this predicate is
evaluated.
Table 62. EXPLAIN_PREDICATE Table (continued). PK means that the column is part of a primary key; FK means
that the column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
PREDICATE_TEXT CLOB(2M) Yes No The text of the predicate as recreated from the internal
| representation of the SQL statement. This will also
| contain the value of the host variable, special register, or
| parameter marker, if used during compilation of the
| statement.
EXPLAIN_STATEMENT table
The EXPLAIN_STATEMENT table contains the text of the SQL statement as it
exists for the different levels of Explain information. The original SQL statement as
entered by the user is stored in this table along with the version used (by the
optimizer) to choose an access plan to satisfy the SQL statement. The latter version
may bear little resemblance to the original as it may have been rewritten and/or
enhanced with additional predicates as determined by the SQL Compiler.
Table 65. EXPLAIN_STATEMENT Table. PK means that the column is part of a primary key; FK means that the
column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
EXPLAIN_REQUESTER VARCHAR(128) No PK, FK Authorization ID of initiator of this Explain request.
EXPLAIN_TIME TIMESTAMP No PK, FK Time of initiation for Explain request.
SOURCE_NAME VARCHAR(128) No PK, FK Name of the package running when the dynamic
statement was explained or name of the source file
when the static SQL was explained.
SOURCE_SCHEMA VARCHAR(128) No PK, FK Schema, or qualifier, of source of Explain request.
SOURCE_VERSION VARCHAR(64) No FK Version of the source of the Explain request.
EXPLAIN_LEVEL CHAR(1) No PK Level of Explain information for which this row is
relevant.
Table 65. EXPLAIN_STATEMENT Table (continued). PK means that the column is part of a primary key; FK means
that the column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
UPDATABLE CHAR(1) No No Indicates if this statement is considered updatable. This
is particularly relevant to SELECT statements which
may be determined to be potentially updatable.
EXPLAIN_STREAM table
The EXPLAIN_STREAM table represents the input and output data streams
between individual operators and data objects. The data objects themselves are
represented in the EXPLAIN_OBJECT table. The operators involved in a data
stream are to be found in the EXPLAIN_OPERATOR table.
Table 66. EXPLAIN_STREAM Table. PK means that the column is part of a primary key; FK means that the column is
part of a foreign key.
Column Name Data Type Nullable? Key? Description
EXPLAIN_REQUESTER VARCHAR(128) No FK Authorization ID of initiator of this Explain request.
EXPLAIN_TIME TIMESTAMP No FK Time of initiation for Explain request.
SOURCE_NAME VARCHAR(128) No FK Name of the package running when the dynamic
statement was explained or name of the source file when
the static SQL was explained.
SOURCE_SCHEMA VARCHAR(128) No FK Schema, or qualifier, of source of Explain request.
SOURCE_VERSION VARCHAR(64) No FK Version of the source of the Explain request.
EXPLAIN_LEVEL CHAR(1) No FK Level of Explain information for which this row is
relevant.
STMTNO INTEGER No FK Statement number within package to which this explain
information is related.
SECTNO INTEGER No FK Section number within package to which this explain
information is related.
STREAM_ID INTEGER No No Unique ID for this data stream within the specified
operator.
SOURCE_TYPE CHAR(1) No No Indicates the source of this data stream:
O Operator
D Data Object
SOURCE_ID SMALLINT No No Unique ID for the operator within this query that is the
source of this data stream. Set to -1 if SOURCE_TYPE is
’D’.
TARGET_TYPE CHAR(1) No No Indicates the target of this data stream:
O Operator
D Data Object
TARGET_ID SMALLINT No No Unique ID for the operator within this query that is the
target of this data stream. Set to -1 if TARGET_TYPE is
’D’.
OBJECT_SCHEMA VARCHAR(128) Yes No Schema to which the affected data object belongs. Set to
null if both SOURCE_TYPE and TARGET_TYPE are ’O’.
OBJECT_NAME VARCHAR(128) Yes No Name of the object that is the subject of data stream. Set
to null if both SOURCE_TYPE and TARGET_TYPE are
’O’.
STREAM_COUNT DOUBLE No No Estimated cardinality of data stream.
COLUMN_COUNT SMALLINT No No Number of columns in data stream.
PREDICATE_ID INTEGER No No If this stream is part of a subquery for a predicate, the
predicate ID will be reflected here, otherwise the column
is set to -1.
Table 66. EXPLAIN_STREAM Table (continued). PK means that the column is part of a primary key; FK means that
the column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
COLUMN_NAMES CLOB(2M) Yes No This column contains the names and ordering
information of the columns involved in this stream.
ADVISE_INDEX table
The ADVISE_INDEX table represents the recommended indexes.
Table 67. ADVISE_INDEX Table. PK means that the column is part of a primary key; FK means that the column is
part of a foreign key.
Column Name Data Type Nullable? Key? Description
EXPLAIN_REQUESTER VARCHAR(128) No No Authorization ID of initiator of this Explain request.
EXPLAIN_TIME TIMESTAMP No No Time of initiation for Explain request.
SOURCE_NAME VARCHAR(128) No No Name of the package running when the dynamic
statement was explained or name of the source file when
static SQL was explained.
SOURCE_SCHEMA VARCHAR(128) No No Schema, or qualifier, of source of Explain request.
SOURCE_VERSION VARCHAR(64) No No Version of the source of the Explain request.
EXPLAIN_LEVEL CHAR(1) No No Level of Explain information for which this row is
relevant.
STMTNO INTEGER No No Statement number within package to which this explain
information is related.
SECTNO INTEGER No No Section number within package to which this explain
information is related.
QUERYNO INTEGER No No Numeric identifier for explained SQL statement. For
dynamic SQL statements (excluding the EXPLAIN SQL
statement) issued through CLP or CLI, the default value
is a sequentially incremented value. Otherwise, the
default value is the value of STMTNO for static SQL
statements and 1 for dynamic SQL statements.
QUERYTAG CHAR(20) No No Identifier tag for each explained SQL statement. For
dynamic SQL statements issued through CLP (excluding
the EXPLAIN SQL statement), the default value is 'CLP'.
For dynamic SQL statements issued through CLI
(excluding the EXPLAIN SQL statement), the default
value is 'CLI'. Otherwise, the default value used is
blanks.
NAME VARCHAR(128) No No Name of the index.
CREATOR VARCHAR(128) No No Qualifier of the index name.
TBNAME VARCHAR(128) No No Name of the table or nickname on which the index is
defined.
TBCREATOR VARCHAR(128) No No Qualifier of the table name.
COLNAMES CLOB(2M) No No List of column names.
UNIQUERULE CHAR(1) No No Unique rule:
D = Duplicates allowed
P = Primary index
U = Unique entries only allowed
COLCOUNT SMALLINT No No Number of columns in the key plus the number of
include columns if any.
IID SMALLINT No No Internal index ID.
NLEAF INTEGER No No Number of leaf pages; −1 if statistics are not gathered.
NLEVELS SMALLINT No No Number of index levels; −1 if statistics are not gathered.
FIRSTKEYCARD BIGINT No No Number of distinct first key values; −1 if statistics are
not gathered.
FULLKEYCARD BIGINT No No Number of distinct full key values; −1 if statistics are not
gathered.
Table 67. ADVISE_INDEX Table (continued). PK means that the column is part of a primary key; FK means that the
column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
CLUSTERRATIO SMALLINT No No Degree of data clustering with the index; −1 if statistics
are not gathered or if detailed index statistics are
gathered (in which case, CLUSTERFACTOR will be used
instead).
CLUSTERFACTOR DOUBLE No No Finer measurement of degree of clustering, or −1 if
detailed index statistics have not been gathered or if the
index is defined on a nickname.
USERDEFINED SMALLINT No No Defined by the user.
SYSTEM_REQUIRED SMALLINT No No 1 if one or the other of the following conditions is
met:
– This index is required for a primary or unique key
constraint, or this index is a dimension block index
or composite block index for a multi-dimensional
clustering (MDC) table.
– This is an index on the (OID) column of a typed
table.
2 if both of the following conditions are met:
– This index is required for a primary or unique key
constraint, or this index is a dimension block index
or composite block index for an MDC table.
– This is an index on the (OID) column of a typed
table.
0 otherwise.
CREATE_TIME TIMESTAMP No No Time when the index was created.
STATS_TIME TIMESTAMP Yes No Last time when any change was made to recorded
statistics for this index. Null if no statistics available.
PAGE_FETCH_PAIRS VARCHAR(254) No No A list of pairs of integers, represented in character form.
Each pair represents the number of pages in a
hypothetical buffer, and the number of page fetches
required to scan the table with this index using that
hypothetical buffer. (Zero-length string if no data
available.)
REMARKS VARCHAR(254) Yes No User-supplied comment, or null.
DEFINER VARCHAR(128) No No User who created the index.
CONVERTED CHAR(1) No No Reserved for future use.
SEQUENTIAL_PAGES INTEGER No No Number of leaf pages located on disk in index key order
with few or no large gaps between them. (−1 if no
statistics are available.)
DENSITY INTEGER No No Ratio of SEQUENTIAL_PAGES to number of pages in
the range of pages occupied by the index, expressed as a
percent (integer between 0 and 100, −1 if no statistics are
available.)
FIRST2KEYCARD BIGINT No No Number of distinct keys using the first two columns of
the index (−1 if no statistics or inapplicable)
FIRST3KEYCARD BIGINT No No Number of distinct keys using the first three columns of
the index (−1 if no statistics or inapplicable)
FIRST4KEYCARD BIGINT No No Number of distinct keys using the first four columns of
the index (−1 if no statistics or inapplicable)
PCTFREE SMALLINT No No Percentage of each index leaf page to be reserved during
initial building of the index. This space is available for
future inserts after the index is built.
Table 67. ADVISE_INDEX Table (continued). PK means that the column is part of a primary key; FK means that the
column is part of a foreign key.
Column Name Data Type Nullable? Key? Description
UNIQUE_COLCOUNT SMALLINT No No The number of columns required for a unique key.
Always <=COLCOUNT. < COLCOUNT only if there a
include columns. −1 if index has no unique key (permits
duplicates)
MINPCTUSED SMALLINT No No If not zero, then online index defragmentation is
enabled, and the value is the threshold of minimum
used space before merging pages.
REVERSE_SCANS CHAR(1) No No Y = Index supports reverse scans
N = Index does not support reverse scans
USE_INDEX CHAR(1) Yes No Y = index recommended or evaluated
N = index not to be recommended
CREATION_TEXT CLOB(2M) No No The SQL statement used to create the index.
PACKED_DESC BLOB(1M) Yes No Internal description of the table.
| ADVISE_INSTANCE table
| The ADVISE_INSTANCE table contains information about db2advis execution,
| including start time. Contains one row for each execution of db2advis. Other
| ADVISE tables have a foreign key (RUN_ID) that links to the START_TIME
| column of the ADVISE_INSTANCE table for rows created during the same Design
| Advisor run.
| Table 68. ADVISE_INSTANCE Table. PK means that the column is part of a primary key; FK means that the column
| is part of a foreign key.
| Column Name Data Type Nullable? Key? Description
| START_TIME TIMESTAMP No PK Time at which db2advis execution begins.
| END_TIME TIMESTAMP No No Time at which db2advis execution ends.
| MODE VARCHAR(4) No No The value that was specified with the -m option on the
| Design Advisor; for example, ’MC’ to specify MQT and
| MDC.
| WKLD_COMPRESSION CHAR(4) No No The workload compression under which the Design
| Advisor was run.
| STATUS CHAR(9) No No The status of a Design Advisor run. Status can be
| ’STARTED’, ’COMPLETED’ (if successful), or an error
| number that is prefixed by ’EI’ for internal errors or ’EX’
| for external errors, in which case the error number
| represents the SQLCODE.
|
|
| ADVISE_MQT table
| The ADVISE_MQT table contains information about materialized query tables
| (MQT) recommended by the Design Advisor.
| Table 69. ADVISE_MQT Table. PK means that the column is part of a primary key; FK means that the column is part
| of a foreign key.
| Column Name Data Type Nullable? Key? Description
| EXPLAIN_REQUESTER VARCHAR(128) No No Authorization ID of initiator of this Explain request.
| EXPLAIN_TIME TIMESTAMP No No Time of initiation for Explain request.
| SOURCE_NAME VARCHAR(128) No No Name of the package running when the dynamic
| statement was explained or name of the source file when
| the static SQL was explained.
| SOURCE_SCHEMA VARCHAR(128) No No Schema, or qualifier, of source of Explain request.
| SOURCE_VERSION VARCHAR(64) No No Version of the source of the Explain request.
| EXPLAIN_LEVEL CHAR(1) No No Level of Explain information for which this row is
| relevant.
| STMTNO INTEGER No No Statement number within package to which this Explain
| information is related.
| SECTNO INTEGER No No Statement number within package to which this Explain
| information is related.
| NAME VARCHAR(128) No No MQT name.
| CREATOR VARCHAR(128) No No MQT creator name.
| IID SMALLINT No No Internal identifier.
| CREATE_TIME TIMESTAMP No No Time at which the MQT was created.
| STATS_TIME TIMESTAMP Yes No Time at which statistics were taken.
| NUMROWS DOUBLE No No The number of estimated rows in the MQT.
| NUMCOLS SMALLINT No No Number of columns defined in the MQT.
| ROWSIZE DOUBLE No No Average length (in bytes) of a row in the MQT.
| BENEFIT FLOAT No No Reserved for future use.
| USE_MQT CHAR(1) Yes No Set to ’Y’ when the MQT is recommended.
| MQT_SOURCE CHAR(1) Yes No Indicates where the MQT candidate was generated. Set
| to ’I’ if the MQT candidate is a refresh-immediate MQT,
| or ’D’ if it can only be created as a full refresh-deferred
| MQT.
| QUERY_TEXT CLOB(2M) No No Contains the query that defines the MQT.
| CREATION_TEXT CLOB(2M) No No Contains the CREATE TABLE DDL for the MQT.
| SAMPLE_TEXT CLOB(2M) No No Contains the sampling query that is used to get detailed
| statistics for the MQT. Only used when detailed statistics
| are required for the Design Advisor. The resulting
| sampled statistics will be shown in this table. If null,
| then no sampling query was created for this MQT.
| COLSTATS CLOB(2M) No No Contains the column statistics for the MQT (if not null).
| These statistics are in XML format and include the
| column name, column cardinality and, optionally, the
| HIGH2KEY and LOW2KEY values.
| EXTRA_INFO BLOB(2M) No No Reserved for miscellaneous output.
| TBSPACE VARCHAR(128) No No The table space that is recommended for the MQT.
| RUN_ID TIMESTAMP Yes FK A value corresponding to the START_TIME of a row in
| the ADVISE_INSTANCE table, linking it to the same
| Design Advisor run.
| REFRESH_TYPE CHAR(1) No No Set to ’I’ for immediate or ’D’ for deferred.
| EXISTS CHAR(1) No No Set to ’Y’ if the MQT exists in the database catalog.
|
|
| ADVISE_PARTITION table
| The ADVISE_PARTITION table contains information about database partitions
| recommended by the Design Advisor, and can only be populated in a partitioned
| database environment.
| Table 70. ADVISE_PARTITION Table. PK means that the column is part of a primary key; FK means that the column
| is part of a foreign key.
| Column Name Data Type Nullable? Key? Description
| EXPLAIN_REQUESTER VARCHAR(128) No No Authorization ID of initiator of this Explain request.
| EXPLAIN_TIME TIMESTAMP No No Time of initiation for Explain request.
| SOURCE_NAME VARCHAR(128) No No Name of the package running when the dynamic
| statement was explained or name of the source file when
| the static SQL was explained.
| SOURCE_SCHEMA VARCHAR(128) No No Schema, or qualifier, of source of Explain request.
| SOURCE_VERSION VARCHAR(64) No No Version of the source of the Explain request.
| EXPLAIN_LEVEL CHAR(1) No No Level of Explain information for which this row is
| relevant.
| STMTNO INTEGER No No Statement number within package to which this Explain
| information is related.
| SECTNO INTEGER No No Statement number within package to which this Explain
| information is related.
| QUERYNO INTEGER No No Numeric identifier for explained SQL statement. For
| dynamic SQL statements (excluding the EXPLAIN SQL
| statement) issued through CLP or CLI, the default value
| is a sequentially incremented value. Otherwise, the
| default value is the value of STMTNO for static SQL
| statements and 1 for dynamic SQL statements.
| QUERYTAG CHAR(20) No No Identifier tag for each explained SQL statement. For
| dynamic SQL statements issued through CLP (excluding
| the EXPLAIN SQL statement), the default value is ’CLP’.
| For dynamic SQL statements issued through CLI
| (excluding the EXPLAIN SQL statement), the default
| value is ’CLI’. Otherwise, the default value used is
| blanks.
| TBNAME VARCHAR(128) Yes No Specifies the table name.
| TBCREATOR VARCHAR(128) Yes No Specifies the table creator name.
| PMID SMALLINT Yes No Specifies the partition map ID.
| TBSPACE VARCHAR(128) Yes No Specifies the table space in which the table resides.
| COLNAMES CLOB(2M) Yes No Specifies partition column names, separated by commas.
| COLCOUNT SMALLINT Yes No Specifies the number of partitioning columns.
| REPLICATE CHAR(1) Yes No Specifies whether or not the partition is replicated.
| COST DOUBLE Yes No Specifies the cost of using the partition.
| USEIT CHAR(1) Yes No Specifies whether or not the partition is used in
| EVALUATE PARTITION mode. A partition is used if
| USEIT is set to ’Y’ or ’y’.
| RUN_ID TIMESTAMP Yes FK A value corresponding to the START_TIME of a row in
| the ADVISE_INSTANCE table, linking it to the same
| Design Advisor run.
|
|
| ADVISE_TABLE table
| The ADVISE_TABLE table stores the data definition language (DDL) for table
| creation, using the final Design Advisor recommendations for materialized query
| tables (MQTs), multidimensional clustered tables (MDCs), and partitioning.
| Table 71. ADVISE_TABLE Table. PK means that the column is part of a primary key; FK means that the column is
| part of a foreign key.
| Column Name Data Type Nullable? Key? Description
| RUN_ID TIMESTAMP Yes FK A value corresponding to the START_TIME of a row in
| the ADVISE_INSTANCE table, linking it to the same
| Design Advisor run.
| TABLE_NAME VARCHAR(128) No No Name of the table.
| TABLE_SCHEMA VARCHAR(128) No No Name of the table creator.
| TABLESPACE VARCHAR(128) No No The table space in which the table is to be created.
| SELECTION_FLAG VARCHAR(4) No No Indicates the recommendation type. Valid values are ’M’
| for MQT, ’P’ for partitioning, and ’C’ for MDC. This
| field can include any subset of these values. For
| example, ’MC’ indicates that the table is recommended
| as an MQT and an MDC table.
| TABLE_EXISTS CHAR(1) No No Set to ’Y’ if the table exists in the database catalog.
| USE_TABLE CHAR(1) No No Set to ’Y’ if the table has recommendations from the
| Design Advisor.
| GEN_COLUMNS CLOB(2M) No No Contains a generated columns string if this row includes
| an MDC recommendation that requires generated
| columns in the create table DDL.
| ORGANIZE_BY CLOB(2M) No No For MDC recommendations, contains the ORGANIZE
| BY clause of the create table DDL.
| CREATION_TEXT CLOB(2M) No No Contains the create table DDL.
| ALTER_COMMAND CLOB(2M) No No Contains an ALTER TABLE statement for the table.
|
|
ADVISE_WORKLOAD table
The ADVISE_WORKLOAD table represents the statement that makes up the
workload.
Table 72. ADVISE_WORKLOAD Table. PK means that the column is part of a primary key; FK means that the column
is part of a foreign key.
Column Name Data Type Nullable? Key? Description
WORKLOAD_NAME CHAR(128) No No Name of the collection of SQL statements (workload)
that this statments belongs to.
STATEMENT_NO INTEGER No No Statement number within the workload to which this
explain information is related.
STATEMENT_TEXT CLOB(1M) No No Content of the SQL statement.
STATEMENT_TAG VARCHAR(256) No No Identifier tag for each explained SQL statement.
FREQUENCY INTEGER No No The number of times this statement appears within the
workload.
IMPORTANCE DOUBLE No No Importance of the statement.
WEIGHT DOUBLE No No Priority of the statement.
COST_BEFORE DOUBLE Yes No The cost (in timerons) of the query if the recommended
indexes are not created.
COST_AFTER DOUBLE Yes No The cost (in timerons) of the query if the recommended
indexes are created.
COMPILABLE CHAR(17) Yes No Indicates any query compile errors that occured while
trying to prepare the statement. If this column is NULL
or does not start with SQLCA, the SQL query could be
compiled by db2advis. If a compile error is found by
db2advis or the Design Advisor, the COMPILABLE
column value consists of an 8-character [Link]
field, followed by a colon (:) and an 8-character
[Link] field, which is the return code for the
SQL statement.
To fully use the output of db2expln, and dynexpln you must understand:
v The different SQL statements supported and the terminology related to those
statements (such as predicates in a SELECT statement)
v The purpose of a package (access plan)
v The purpose and contents of the system catalog tables
v General application tuning concepts
The topics in this section provide information about db2expln and dynexpln.
The dynexpln tool can also be used to describe the access plan selected for dynamic
statements. It creates a static package for the statements and then uses the
db2expln tool to describe them. However, because the dynamic SQL can be
examined by db2expln this utility is retained only for backward compatibility.
The explain tools (db2expln and dynexpln) are located in the bin subdirectory of
your instance sqllib directory. If db2expln and dynexpln are not in your current
directory, they must be in a directory that appears in your PATH environment
variable.
The db2expln program connects and uses the [Link], [Link], and
[Link] files to bind itself to a database the first time the database is
accessed.
To run db2expln, you must have the SELECT privilege on the system catalog views
as well as the EXECUTE privilege for the db2expln, db2exsrv, and db2exdyn
packages. To run dynexpln, you must have BINDADD authority for the database,
and the schema you are using to connect to the database must exist or you must
have the IMPLICIT_SCHEMA authority for the database. To explain dynamic SQL
using either db2expln or dynexpln, you must also have any privileges needed for
the SQL statements being explained. (Note that if you have SYSADM or DBADM
authority, you will automatically have all these authorization levels.)
db2expln
The following sections describe the syntax and parameters for db2expln and
provide usage notes.
db2expln
connection-options output-options
package-options dynamic-options explain-options
-help
connection-options:
-database database-name
-user user-id password
output-options:
package-options:
-version version-identifier -escape escape-character
-noupper -section section-number
dynamic-options:
-statement sql-statement -stmtfile sql-statement-file
-terminator termination-character -noenv
explain-options:
-graph -opids
Command parameters:
The options may be specified in any order.
connection-options:
These options specify the database to connect to and any options necessary to
make the connection. The connection options are required except when the -help
option is specified.
-database database-name
The name of the database that contains the packages to be explained.
For backward compatibility, you can use -d instead of -database.
-user user-id password
The authorization ID and password to use when establishing the database
connection. Both user-id and password must be valid according to DB2®
naming conventions and must be recognized by the database.
For backward compatibility, you can use -u instead of -user.
output-options:
These options specify where the db2expln output should be directed. Except when
the -help option is specified, you must specify at least one output option. If you
specify both options, output is sent to a file as well as to the terminal.
-output output-file
The output of db2expln is written to the file that you specify.
For backward compatibility, you can use -o instead of -output.
-terminal
The db2expln output is directed to the terminal.
For backward compatibility, you can use -t instead of -terminal.
package-options:
These options specify one or more packages and sections to be explained. Only
static SQL in the packages and sections is explained.
Note: As in a LIKE predicate, you can use the pattern matching characters, which
are percent sign (%) and underscore (_), to specify the schema-name,
package-name, and version-identifier.
-schema schema-name
The schema of the package or packages to be explained.
For backward compatibility, you can use -c instead of -schema.
-package package-name
The name of the package or packages to be explained.
For backward compatibility, you can use -p instead of -package.
-version version-identifier
The version identifier of the package or packages to be explained. The
default version is the empty string.
-escape escape-character
The character, escape-character to be used as the escape character for pattern
matching in the schema-name, package-name, and version-identifier.
For example, the db2expln command to explain the package
[Link]% is as follows:
db2expln -schema TESTID -package CALC% ....
However, this command would also explain any other plans that start with
CALC. To explain only the [Link]% package, you must use an
escape character. If you specify the exclamation point (!) as the escape
character, you can change the command to read: db2expln -schema TESTID
-escape ! -package CALC!% ... . Then the ! character is used as an escape
character and thus !% is interpreted as the % character and not as the
″match anything″ pattern. There is no default escape character.
dynamic-options:
These statements make it possible to alter the plan chosen for subsequent
dynamic SQL statements processed by db2expln.
If you specify -noenv, then these statement are explained, but not
executed.
explain-options:
-help Shows the help text for db2expln. If this option is specified no packages are
explained.
Most of the command line is processed in the db2exsrv stored procedure. To
get help on all the available options, it is necessary to provide
connection-options along with -help. For example, use:
db2expln -help -database SAMPLE
Usage notes:
Unless you specify the -help option, you must specify either package-options or
dynamic-options. You can explain both packages and dynamic SQL with a single
invocation of db2expln.
Some of the option flags above might have special meaning to your operating
system and, as a result, might not be interpreted correctly in the db2expln
command line. However, you might be able to enter these characters by preceding
them with an operating system escape character. For more information, see your
operating system documentation. Make sure that you do not inadvertently specify
the operating system escape character as the db2expln escape character.
Help and initial status messages, produced by db2expln, are written to standard
output. All prompts and other status messages produced by the explain tool are
written to standard error. Explain text is written to standard output or to a file
depending on the output option chosen.
Examples:
To explain multiple plans with one invocation of db2expln, use the -package,
-schema, and -version option and specify string constants for packages and
creators with LIKE patterns. That is, the underscore (_) may be used to represent a
single character, and the percent sign (%) may be used to represent the occurrence
of zero or more characters.
To explain all sections for all packages in a database named SAMPLE, with the
results being written to the file [Link] , enter
db2expln -database SAMPLE -schema % -package % -output [Link]
As another example, suppose a user has a CLP script file called ″statements.db2″
and wants to explain the statements in the file. The file contains the following
statements:
SET PATH=SYSIBM, SYSFUN, DEPT01, DEPT93@
SELECT EMPNO, TITLE(JOBID) FROM EMPLOYEE@
Related concepts:
v “SQL explain tools” on page 551
v “Description of db2expln and dynexpln output” on page 558
v “Examples of db2expln and dynexpln output” on page 576
Each sub-statement within a compound SQL statement may have its own section,
which can be explained by db2expln.
dynexpln
The dynexpln tool is still available for backward compatibility. However, you can
use the dynamic-options of db2expln to perform all of the functions of dynexpln.
When you use the dynamic-options of db2expln, the statement is prepared as true
dynamic SQL and the generated plan is explained from the SQL cache. This
explain-output method provides more accurate access plans than dynexpln, which
prepares the statement as static SQL. It also allows the use of features available
only in dynamic SQL, such as parameter markers.
Related concepts:
v “SQL explain tools” on page 551
v “Examples of db2expln and dynexpln output” on page 576
The steps of an access plan, or section, are presented in the order that the database
manager executes them. Each major step is shown as a left-justified heading with
information about that step indented below it. Indentation bars appear in the left
margin of the explain output for the access plan. These bars also mark the scope of
the operation. Operations at a lower level of indentation, farther to the right, in the
same operation are processed before returning to the previous level of indentation.
Remember that the access plan chosen was based on an augmented version of the
original SQL statement that is shown in the output. For example, the original
statement may cause triggers and constraints to be activated. In addition, the query
rewrite component of the SQL compiler might rewrite the SQL statement to an
equivalent but more efficient format. All of these factors are included in the
information that the optimizer uses when it determines the most efficient plan to
satisfy the statement. Thus, the access plan shown in the explain output may differ
substantially from the access plan that you might expect for the original SQL
statement. The SQL Explain facility, which includes the explain tables, SET
CURRENT EXPLAIN mode, and Visual Explain, shows the actual SQL statement
used for optimization in the form of an SQL-like statement which is created by
reverse-translating the internal representation of the query.
When you compare output from db2expln or dynexpln to the output of the Explain
facility, the operator ID option (-opids) can be very useful. Each time db2expln or
dynexpln starts processing a new operator from the Explain facility, the operator
ID number is printed to the left of the explained plan. The operator IDs can be
used to match up the steps in the different representations of the access plan. Note
that there is not always a one-to-one correspondence between the operators in the
Explain facility output and the operations shown by db2expln and dynexpln.
Related concepts:
v “dynexpln” on page 558
v “Table access information” on page 559
v “Temporary table information” on page 564
v “Join information” on page 566
v “Data stream information” on page 568
v “Insert, update, and delete information” on page 568
v “Block and row identifier preparation information” on page 569
v “Aggregation information” on page 570
v “Parallel processing information” on page 571
v “Federated query information” on page 573
v “Miscellaneous information” on page 574
Related reference:
v “db2expln - SQL Explain” on page 552
where:
– [Link] is the fully-qualified name of the table being accessed
– ID is the corresponding TABLESPACEID and TABLEID from the
[Link] catalog for the table
v Access Hierarchy Table Name:
Access Hierarchy Table Name = [Link] ID = ts,n
where:
– [Link] is the fully-qualified name of the table being accessed
– ID is the corresponding TABLESPACEID and TABLEID from the
[Link] catalog for the table
v Access Materialized Query Table Name:
Access Materialized Query Table Name = [Link] ID = ts,n
where:
– [Link] is the fully-qualified name of the table being accessed
– ID is the corresponding TABLESPACEID and TABLEID from the
[Link] catalog for the table
2. Temporary tables of two types:
v Access Temporary Table ID:
Access Temp Table ID = tn
where:
– ID is the corresponding identifier assigned by db2expln
v Access Declared Global Temporary Table ID:
Access Global Temp Table ID = ts,tn
where:
– ID is the corresponding TABLESPACEID from the [Link]
catalog for the table (ts); and the corresponding identifier assigned by
db2expln (tn)
Number of Columns
The following statement indicates the number of columns being used from each
row of the table:
#Columns = n
Block Access
The following statement indicates that the table has one or more dimension block
indexes defined on it:
Clustered by Dimension for Block Index Access
If this text is not shown, the table was created without the DIMENSION clause.
Parallel Scan
The following statement indicates that the database manager will use several
subagents to read from the table in parallel:
Parallel Scan
If this text is not shown, the table will only be read from by one agent (or
subagent).
Scan Direction
The following statement indicates that the database manager will read rows in a
reverse order:
Scan Direction = Reverse
If this text is not shown, the scan direction is forward, which is the default.
One of the following statements will be displayed, indicating how the qualifying
rows in the table are being accessed:
v The Relation Scan statement indicates that the table is being sequentially
scanned to find the qualifying rows.
– The following statement indicates that no prefetching of data will be done:
Relation Scan
| Prefetch: None
– The following statement indicates that the optimizer has predetermined the
number of pages that will be prefetched:
Relation Scan
| Prefetch: n Pages
– The following statement indicates that data should be prefetched:
Relation Scan
| Prefetch: Eligible
– The following statement indicates that the qualifying rows are being
identified and accessed through an index:
Index Scan: Name = [Link] ID = xx
| Index type
| Index Columns:
where:
- [Link] is the fully-qualified name of the index being scanned
- ID is the corresponding IID column in the [Link] catalog view.
- Index type is one of:
Regular Index (Not Clustered)
Regular Index (Clustered)
Dimension Block Index
Composite Dimension Block Index
This will be followed by one row for each column in the index. Each
column in the index will be listed in one of the following forms:
n: column_name (Ascending)
n: column_name (Descending)
n: column_name (Include Column)
The following statements are provided to clarify the type of index scan:
- The range delimiting predicates for the index are shown by:
#Key Columns = n
| Start Key: xxxxx
| Stop Key: xxxxx
Where xxxxx is one of:
v Start of Index
v End of Index
v Inclusive Value: or Exclusive Value:
An inclusive key value will be included in the index scan. An exclusive
key value will not be included in the scan. The value for the key will be
given by one of the following rows for each part of the key:
n: ’string’
n: nnn
n: yyyy-mm-dd
n: hh:mm:ss
n: yyyy-mm-dd hh:mm:[Link]
n: NULL
n: ?
If a literal string is shown, only the first 20 characters are displayed. If
the string is longer than 20 characters, this will be shown by ... at the
end of the string. Some keys cannot be determined until the section is
executed. This is shown by a ? as the value.
- Index-Only Access
If all the needed columns can be obtained from the index key, this
statement will appear and no table data will be accessed.
- The following statement indicates that no prefetching of index pages will
be done:
Index Prefetch: None
- The following statement indicates that index pages should be prefetched:
Index Prefetch: Eligible
- The following statement indicates that no prefetching of data pages will be
done:
Data Prefetch: None
- The following statement indicates that data pages should be prefetched:
Data Prefetch: Eligible
- If there are predicates that can be passed to the Index Manager to help
qualify index entries, the following statement is used to show the number
of predicates:
Sargable Index Predicate(s)
| #Predicates = n
– If the qualifying rows are being accessed by using row IDs (RIDs) that were
prepared earlier in the access plan, it will be indicated with the statement:
Fetch Direct Using Row IDs
If the table has one or more block indexes defined for it, then rows may be
accessed by either block or row IDs. This is indicated by:
Lock Intents
For each table access, the type of lock that will be acquired at the table and row
levels is shown with the following statement:
Lock Intents
| Table: xxxx
| Row : xxxx
Predicates
There are two statements that provide information about the predicates used in an
access plan:
1. The following statement indicates the number of predicates that will be
evaluated for each block of data retrieved from a blocked index.
Block Predicates(s)
| #Predicates = n
2. The following statement indicates the number of predicates that will be
evaluated while the data is being accessed. The count of predicates does not
include push-down operations such as aggregation or sort.
Sargable Predicate(s)
| #Predicates = n
3. The following statement indicates the number of predicates that will be
evaluated once the data has been returned:
Residual Predicate(s)
| #Predicates = n
The number of predicates shown in the above statements may not reflect the
number of predicates provided in the SQL statement because predicates can be:
v Applied more than once within the same query
v Transformed and extended with the addition of implicit predicates during the
query optimization process
v Transformed and condensed into fewer predicates during the query optimization
process.
v The following statement indicates that some or all of the rows read from the
temporary table will be cached outside the buffer pool if sufficient sortheap
memory is available:
Keep Rows In Private Memory
v If the table has the volatile cardinality attribute set, it will be indicated by:
Volatile Cardinality
Related concepts:
v “Description of db2expln and dynexpln output” on page 558
v “Examples of db2expln and dynexpln output” on page 576
If a temporary table needs to be created, then one of two possible statements may
appear. These statements indicate that a temporary table is to be created and rows
inserted into it. The ID is an identifier assigned by db2expln for convenience when
referring to the temporary table. This ID is prefixed with the letter ’t’ to indicate
that the table is a temporary table.
v The following statement indicates an ordinary temporary table will be created:
Insert Into Temp Table ID = tn
v The following statement indicates an ordinary temporary table will be created by
multiple subagents in parallel:
Insert Into Shared Temp Table ID = tn
v The following statement indicates a sorted temporary table will be created:
Insert Into Sorted Temp Table ID = tn
v The following statement indicates a sorted temporary table will be created by
multiple subagents in parallel:
Insert Into Sorted Shared Temp Table ID = tn
v The following statement indicates a declared global temporary table will be
created:
Insert Into Global Temp Table ID = ts,tn
v The following statement indicates a declared global temporary table will be
created by multiple subagents in parallel:
Insert Into Shared Global Temp Table ID = ts,tn
v The following statement indicates a sorted declared global temporary table will
be created:
Insert Into Sorted Global Temp Table ID = ts,tn
v The following statement indicates a sorted declared global temporary table will
be created by multiple subagents in parallel:
Insert Into Sorted Shared Global Temp Table ID = ts,tn
which indicates how many columns are in each row being inserted into the
temporary table.
A number of additional statements may follow the original creation statement for a
sorted temporary table:
v The following statement indicates the number of key columns used in the sort:
#Sort Key Columns = n
For each column in the sort key, one of the following lines will be displayed:
Key n: column_name (Ascending)
Key n: column_name (Descending)
Key n: (Ascending)
Key n: (Descending)
v The following statements provide estimates of the number of rows and the row
size so that the optimal sort heap can be allocated at run time.
Sortheap Allocation Parameters:
| #Rows = n
| Row Width = n
v If only the first rows of the sorted result are needed, the following is displayed:
Sort Limited To Estimated Row Count
v For sorts in a symmetric multiprocessor (SMP) environment, the type of sort to
be performed is indicated by one of the following statements:
Use Partitioned Sort
Use Shared Sort
Use Replicated Sort
Use Round-Robin Sort
v The following statements indicate whether or not the result from the sort will be
left in the sort heap:
Piped
and
Not Piped
If a piped sort is indicated, the database manager will keep the sorted output in
memory, rather than placing the sorted result in another temporary table.
v The following statement indicates that duplicate values will be removed during
the sort:
Duplicate Elimination
v If aggregation is being performed in the sort, it will be indicated by one of the
following statements:
Partial Aggregation
Intermediate Aggregation
Buffered Partial Aggregation
Buffered Intermediate Aggregation
Table Functions
Table functions are user-defined functions (UDFs) that return data to the statement
in the form of a table. Table functions are indicated by:
Access User Defined Table Function
| Name = [Link]
| Specific Name = specificname
| SQL Access Level = accesslevel
| Language = lang
| Parameter Style = parmstyle
| Fenced Not Deterministic
| Called on NULL Input Disallow Parallel
| Not Federated Not Threadsafe
The specific name uniquely identifies the table function invoked. The remaining
rows detail the atributes of the function.
Related concepts:
v “Description of db2expln and dynexpln output” on page 558
v “Examples of db2expln and dynexpln output” on page 576
Join information
There are three types of joins:
v Hash join
v Merge join
v Nested loop join.
When the time comes in the execution of a section for a join to be performed, one
of the following statements is displayed:
Hash Join
Merge Join
Nested Loop Join
It is possible for a left outer join to be performed. A left outer join is indicated by
one of the following statements:
Left Outer Hash Join
Left Outer Merge Join
Left Outer Nested Loop Join
For merge and nested loop joins, the outer table of the join will be the table
referenced in the previous access statement shown in the output. The inner table of
the join will be the table referenced in the access statement that is contained within
the scope of the join statement. For hash joins, the access statements are reversed
with the outer table contained within the scope of the join and the inner table
appearing before the join.
For a hash or merge join, the following additional statements may appear:
v In some circumstances, a join simply needs to determine if any row in the inner
table matches the current row in the outer. This is indicated with the statement:
Early Out: Single Match Per Outer Row
v It is possible to apply predicates after the join has completed. The number of
predicates being applied will be indicated as follows:
Residual Predicate(s)
| #Predicates = n
For a nested loop join, the following additional statement may appear immediately
after the join statement:
Piped Inner
This statement indicates that the inner table of the join is the result of another
series of operations. This is also referred to as a composite inner.
If a join involves more than two tables, the explain steps should be read from top
to bottom. For example, suppose the explain output has the following flow:
Access ..... W
Join
| Access ..... X
Join
| Access ..... Y
Join
| Access ..... Z
Related concepts:
v “Description of db2expln and dynexpln output” on page 558
v “Examples of db2expln and dynexpln output” on page 576
where n is a unique identifier assigned by db2expln for ease of reference. The end
of a data stream is indicated by:
End of Data Stream n
All operations between these statements are considered part of the same data
stream.
A data stream has a number of characteristics and one or more statements can
follow the initial data stream statement to describe these characteristics:
v If the operation of the data stream depends on a value generated earlier in the
access plan, the data stream is marked with:
Correlated
v Similar to a sorted temporary table, the following statements indicate whether or
not the results of the data stream will be kept in memory:
Piped
and
Not Piped
As was the case with temporary tables, a piped data stream may be written to
disk, if insufficient memory exists at execution time. The access plan will
provide for both possibilities.
v The following statement indicates that only a single record is required from this
data stream:
Single Record
When a data stream is accessed, the following statement will appear in the output:
Access Data Stream n
Related concepts:
v “Description of db2expln and dynexpln output” on page 558
v “Examples of db2expln and dynexpln output” on page 576
Related concepts:
v “Description of db2expln and dynexpln output” on page 558
v “Examples of db2expln and dynexpln output” on page 576
Index ORing refers to the technique of making more than one index access and
combining the results to include the distinct IDs that appear in any of the
indexes accessed. The optimizer will consider index ORing when predicates are
connected by OR keywords or there is an IN predicate. The index accesses can
be on the same index or different indexes.
v Another use of ID preparation is to prepare the input data to be used during list
prefetch, as indicated by either of the following:
List Prefetch Preparation
Block List Prefetch RID Preparation
v Index ANDing refers to the technique of making more than one index access and
combining the results to include IDs that appear in all of the indexes accessed.
Index ANDing processing is started with either of these statements:
Index ANDing
Block Index ANDing
If the optimizer has estimated the size of the result set, the estimate is shown
with the following statement:
Optimizer Estimate of Set Size: n
Index ANDing filter operations process IDs and use bit filter techniques to
determine the IDs which appear in every index accessed. The following
statements indicate that IDs are being processed for index ANDing:
Index ANDing Bitmap Build Using Row IDs
Index ANDing Bitmap Probe Using Row IDs
Index ANDing Bitmap Build and Probe Using Row IDs
Block Index ANDing Bitmap Build Using Block IDs
Block Index ANDing Bitmap Build and Probe Using Block IDs
Block Index ANDing Bitmap Build and Probe Using Row IDs
Block Index ANDing Bitmap Probe Using Block IDs and Build Using Row IDs
Block Index ANDing Bitmap Probe Using Block IDs
Block Index ANDing Bitmap Probe Using Row IDs
If the optimizer has estimated the size of the result set for a bitmap, the estimate
is shown with the following statement:
Optimizer Estimate of Set Size: n
Related concepts:
v “Description of db2expln and dynexpln output” on page 558
v “Examples of db2expln and dynexpln output” on page 576
Aggregation information
Aggregation is performed on those rows meeting the specified criteria, if any,
provided by the SQL statement predicates. If some sort of aggregate function is to
be done, one of the following statements appears:
Aggregation
Predicate Aggregation
Partial Aggregation
Partial Predicate Aggregation
Intermediate Aggregation
Intermediate Predicate Aggregation
Final Aggregation
Final Predicate Aggregation
Predicate aggregation states that the aggregation operation has been pushed-down
to be processed as a predicate when the data is actually accessed.
Beneath either of the above aggregation statements will be a indication of the type
of aggregate function being performed:
Group By
Column Function(s)
Single Record
The specific column function can be derived from the original SQL statement. A
single record is fetched from an index to satisfy a MIN or MAX operation.
Related concepts:
v “Description of db2expln and dynexpln output” on page 558
v “Examples of db2expln and dynexpln output” on page 576
– The following statements indicate that data is being inserted into a table
queue:
Insert Into Synchronous Table Queue ID = qn
Insert Into Asynchronous Table Queue ID = qn
Insert Into Synchronous Local Table Queue ID = qn
Insert Into Asynchronous Local Table Queue ID = qn
– For database partition table queues, the destination for rows inserted into the
table queue is described by one of the following:
All rows are sent to the coordinator node:
Broadcast to Coordinator Node
All rows are sent to every database partition where the given subsection is
running:
Broadcast to All Nodes of Subsection n
Each row is sent to a database partition based on the values in the row:
Hash to Specific Node
These messages are followed by an indication of the number of keys used for
the sort operation.
#Key Columns = n
For each column in the sort key, one of the following is displayed:
Key n: (Ascending)
Key n: (Descending)
– If predicates will be applied to rows by the receiving end of the table queue,
the following message is shown:
Residual Predicate(s)
| #Predicates = n
v Some subsections in a partitioned database environment explicitly loop back to
the start of the subsection with the statement:
Jump Back to Start of Subsection
Related concepts:
v “Description of db2expln and dynexpln output” on page 558
v “Examples of db2expln and dynexpln output” on page 576
An insert, update, or delete operation that occurs at a data source will be indicated
by the appropriate message:
Ship Distributed Insert #n
Ship Distributed Update #n
Ship Distributed Delete #n
If a table is being explicitly locked at a data source, this will be indicated with the
statement:
Ship Distributed Lock Table #n
DDL statements against a data source are split into two parts. The part invoked at
the data source is indicated by:
Ship Distributed DDL Statement #n
If the federated server is a partitioned database, then part of the DDL statement
must be run at he catalog node. This is indicated by:
Distributed DDL Statement #n Completion
The detail for each distributed substatement is provided separately. The options for
distributed statements are described below:
v The data source for the subquery is shown by one of the following:
Server: server_name (type, version)
Server: server_name (type)
Server: server_name
v If the data source is relational, the SQL for the substatement is displayed as:
SQL Statement:
statement
Non-relational data sources are indicated with:
Non-Relational Data Source
v The nicknames referenced in the substatement are listed as follows:
Nicknames Referenced:
[Link] ID = n
If the data source is relational, the base table for the nickname is shows as:
Base = [Link]
If the data source is non-relational, the source file for the nickname is shown as:
Source File = filename
v If values are passed from the federated server to the data source before
executing the substatement, the number of values will be shown by:
#Input Columns: n
v If values are passed from the data source to the federated server after executing
the substatement, the number of values will be shown by:
#Output Columns: n
Related concepts:
v “Guidelines for analyzing where a federated query is evaluated” on page 182
v “Description of db2expln and dynexpln output” on page 558
Miscellaneous information
v Sections for data definition language statements will be indicated in the output
with the following:
DDL Statement
If the position operation is against a federated data source, then the statement is:
This statement would appear for any SQL statement that uses the WHERE
CURRENT OF syntax.
v The following statement will appear if there are predicates that must be applied
to the result but that could not be applied as part of another operation:
Residual Predicate Application
| #Predicates = n
v The following statement will appear if there is a UNION operator in the SQL
statement:
UNION
v The following statement will appear if there is an operation in the access plan,
whose sole purpose is to produce row values for use by subsequent operations:
Table Constructor
| n-Row(s)
Table constructors can be used for transforming values in a set into a series of
rows that are then passed to subsequent operations. When a table constructor is
prompted for the next row, the following statement will appear:
Access Table Constructor
v The following statement will appear if there is an operation which is only
processed under certain conditions:
Conditional Evaluation
| Condition #n:
| #Predicates = n
| Action #n:
Related concepts:
v “Description of db2expln and dynexpln output” on page 558
v “Examples of db2expln and dynexpln output” on page 576
Related concepts:
v “dynexpln” on page 558
v “Example one: no parallelism” on page 576
v “Example two: single-partition plan with intra-partition parallelism” on page 578
v “Example three: multipartition plan with inter-partition parallelism” on page 579
v “Example four: multipartition plan with inter-partition and intra-partition
parallelism” on page 582
v “Example five: federated database plan” on page 584
Related reference:
v “db2expln - SQL Explain” on page 552
Partition Parallel = No
Intra-Partition Parallel = No
SQL Statement:
DECLARE EMPCUR CURSOR
FOR
SELECT [Link], [Link], [Link], [Link], [Link]
FROM employee AS e, department AS d, project AS p
WHERE [Link] = [Link] AND [Link] = [Link]
End of section
Optimizer Plan:
RETURN
( 1)
|
HSJOIN
( 2)
/ \
HSJOIN TBSCAN
( 3) ( 6)
/ \ |
TBSCAN TBSCAN Table:
( 4) ( 5) DOOLE
| | EMPLOYEE
Table: Table:
DOOLE DOOLE
DEPARTMENT PROJECT
The first part of the plan accesses the DEPARTMENT and PROJECT tables and
uses a hash join to join them. The result of this join is joined to the EMPLOYEE
table. The resulting rows are returned to the application.
Partition Parallel = No
Intra-Partition Parallel = Yes (Bind Degree = 4)
SQL Statement:
DECLARE EMPCUR CURSOR
FOR
SELECT [Link], [Link], [Link], [Link], [Link]
FROM employee AS e, department AS d, project AS p
WHERE [Link] = [Link] AND [Link] = [Link]
End of section
Optimizer Plan:
RETURN
( 1)
|
LTQ
( 2)
|
HSJOIN
( 3)
/ \
HSJOIN TBSCAN
( 4) ( 7)
/ \ |
TBSCAN TBSCAN Table:
( 5) ( 6) DOOLE
| | EMPLOYEE
Table: Table:
DOOLE DOOLE
DEPARTMENT PROJECT
This plan is almost identical to the plan in the first example. The main differences
are the creation of four subagents when the plan first starts and the table queue at
the end of the plan to gather the results of each of subagent’s work before
returning them to the application.
SQL Statement:
DECLARE EMPCUR CURSOR
FOR
SELECT [Link], [Link], [Link], [Link], [Link]
FROM employee AS e, department AS d, project AS p
WHERE [Link] = [Link] AND [Link] = [Link]
Coordinator Subsection:
(-----) Distribute Subsection #2
| Broadcast to Node List
| | Nodes = 10, 33, 55
(-----) Distribute Subsection #3
| Broadcast to Node List
| | Nodes = 10, 33, 55
(-----) Distribute Subsection #1
| Broadcast to Node List
| | Nodes = 10, 33, 55
( 2) Access Table Queue ID = q1 #Columns = 5
( 1) Return Data to Application
| #Columns = 5
Subsection #1:
( 8) Access Table Queue ID = q2 #Columns = 2
( 3) Hash Join
| Estimated Build Size: 5737
| Estimated Probe Size: 8015
( 6) | Access Table Queue ID = q3 #Columns = 3
( 4) | Hash Join
| | Estimated Build Size: 5333
| | Estimated Probe Size: 6421
( 5) | | Access Table Name = [Link] ID = 2,4
| | | #Columns = 3
| | | Relation Scan
| | | | Prefetch: Eligible
| | | Lock Intents
| | | | Table: Intent Share
| | | | Row : Next Key Share
( 5) | | | Process Probe Table for Hash Join
( 2) Insert Into Asynchronous Table Queue ID = q1
| Broadcast to Coordinator Node
| Rows Can Overflow to Temporary Table
Subsection #2:
( 9) Access Table Name = [Link] ID = 2,7
| #Columns = 2
| Relation Scan
| | Prefetch: Eligible
| Lock Intents
| | Table: Intent Share
| | Row : Next Key Share
( 9) | Insert Into Asynchronous Table Queue ID = q2
| | Hash to Specific Node
| | Rows Can Overflow to Temporary Tables
( 8) Insert Into Asynchronous Table Queue Completion ID = q2
Subsection #3:
( 7) Access Table Name = [Link] ID = 2,5
| #Columns = 3
| Relation Scan
| | Prefetch: Eligible
| Lock Intents
| | Table: Intent Share
| | Row : Next Key Share
( 7) | Insert Into Asynchronous Table Queue ID = q3
| | Hash to Specific Node
| | Rows Can Overflow to Temporary Tables
( 6) Insert Into Asynchronous Table Queue Completion ID = q3
End of section
Optimizer Plan:
RETURN
( 1)
|
BTQ
( 2)
|
HSJOIN
( 3)
/ \
HSJOIN DTQ
( 4) ( 8)
/ \ |
TBSCAN DTQ TBSCAN
( 5) ( 6) ( 9)
| | |
Table: TBSCAN Table:
DOOLE ( 7) DOOLE
DEPARTMENT | PROJECT
Table:
DOOLE
EMPLOYEE
This plan has all the same pieces as the plan in the first example, but the section
has been broken into four subsections. The subsections have the following tasks:
v Coordinator Subsection. This subsection coordinates the other subsections. In
this plan, it causes the other subsections to be distributed and then uses a table
queue to gather the results to be returned to the application.
v Subsection #1. This subsection scans table queue q2 and uses a hash join to join
it with the data from table queue q3. A second hash join then adds in the data
from the DEPARTMENT table. The joined rows are then sent to the coordinator
subsection using table queue q1.
v Subsection #2. This subsection scans the PROJECT table and hashes to a specific
node with the results. These results are read by Subsection #1.
v Subsection #3. This subsection scans the EMPLOYEE table and hashes to a
specific node with the results. These results are read by Subsection #1.
SQL Statement:
DECLARE EMPCUR CURSOR
FOR
SELECT [Link], [Link], [Link], [Link], [Link]
FROM employee AS e, department AS d, project AS p
WHERE [Link] = [Link] AND [Link] = [Link]
Coordinator Subsection:
(-----) Distribute Subsection #2
| Broadcast to Node List
| | Nodes = 10, 33, 55
(-----) Distribute Subsection #3
| Broadcast to Node List
| | Nodes = 10, 33, 55
(-----) Distribute Subsection #1
| Broadcast to Node List
| | Nodes = 10, 33, 55
( 2) Access Table Queue ID = q1 #Columns = 5
( 1) Return Data to Application
| #Columns = 5
Subsection #1:
( 3) Process Using 4 Subagents
( 10) | Access Table Queue ID = q3 #Columns = 2
( 4) | Hash Join
| | Estimated Build Size: 5737
Subsection #2:
( 11) Process Using 4 Subagents
( 12) | Access Table Name = [Link] ID = 2,7
| | #Columns = 2
| | Parallel Scan
| | Relation Scan
| | | Prefetch: Eligible
| | Lock Intents
| | | Table: Intent Share
| | | Row : Next Key Share
( 11) | Insert Into Asynchronous Local Table Queue ID = q4
( 11) Access Local Table Queue ID = q4 #Columns = 2
( 10) Insert Into Asynchronous Table Queue ID = q3
| Hash to Specific Node
| Rows Can Overflow to Temporary Tables
Subsection #3:
( 8) Process Using 4 Subagents
( 9) | Access Table Name = [Link] ID = 2,5
| | #Columns = 3
| | Parallel Scan
| | Relation Scan
| | | Prefetch: Eligible
| | Lock Intents
| | | Table: Intent Share
| | | Row : Next Key Share
( 8) | Insert Into Asynchronous Local Table Queue ID = q6
( 8) Access Local Table Queue ID = q6 #Columns = 3
( 7) Insert Into Asynchronous Table Queue ID = q5
| Hash to Specific Node
| Rows Can Overflow to Temporary Tables
End of section
Optimizer Plan:
RETURN
( 1)
|
BTQ
( 2)
|
LTQ
( 3)
|
HSJOIN
( 4)
/ \
HSJOIN DTQ
( 5) ( 10)
/ \ |
TBSCAN DTQ LTQ
( 6) ( 7) ( 11)
| | |
Table: LTQ TBSCAN
DOOLE ( 8) ( 12)
DEPARTMENT | |
TBSCAN Table:
( 9) DOOLE
| PROJECT
Table:
DOOLE
EMPLOYEE
This plan is similar to that in the third example, except that multiple subagents
execute each subsection. Also, at the end of each subsection, a local table queue
gathers the results from all of the subagents before the qualifying rows are inserted
into the second table queue to be hashed to a specific node.
Partition Parallel = No
Intra-Partition Parallel = No
SQL Statement:
DECLARE EMPCUR CURSOR
FOR
SELECT [Link], [Link], [Link], [Link], [Link]
FROM employee AS e, department AS d, project AS p
WHERE [Link] = [Link] AND [Link] = [Link]
( 2) Hash Join
| Estimated Build Size: 48444
| Estimated Probe Size: 232571
( 6) | Access Table Name = [Link] ID = 2,5
| | #Columns = 3
| | Relation Scan
| | | Prefetch: Eligible
| | Lock Intents
| | | Table: Intent Share
| | | Row : Next Key Share
( 6) | | Process Build Table for Hash Join
( 3) | Hash Join
| | Estimated Build Size: 7111
| | Estimated Probe Size: 64606
( 4) | | Ship Distributed Subquery #1
| | | #Columns = 3
( 1) Return Data to Application
| #Columns = 5
Nicknames Referenced:
[Link] ID = 32768
Base = [Link]
#Output Columns = 3
Nicknames Referenced:
[Link] ID = 32769
Base = [Link]
#Output Columns = 2
End of section
Optimizer Plan:
RETURN
( 1)
|
HSJOIN
( 2)
/ \
HSJOIN SHIP
( 3) ( 7)
/ \ |
SHIP TBSCAN Nickname:
( 4) ( 6) DOOLE
| | PROJECT
Nickname: Table:
DOOLE DOOLE
DEPARTMENT EMPLOYEE
This plan has all the same pieces as the plan in the first example, except that the
data for two of the tables are coming from data sources. The two tables are
accessed through distributed subqueries which, in this case, simply select all the
rows from those tables. Once the data is returned to the federated server, it is
joined to the data from the local table.
To use the tool, you require read access to the explain tables being formatted.
Command syntax:
db2exfmt
-d dbname -e schema -f O
-g
O
T
I
C
-l -n name -s schema -o outfile
-t
-u userID password -w timestamp -# sectnbr -h
Command parameters:
-d dbname
Name of the database containing packages.
-e schema
Explain table schema.
-f Formatting flags. In this release, the only supported value is O (operator
summary).
-g Graph plan. If only -g is specified, a graph, followed by formatted
information for all of the tables, is generated. Otherwise, any combination
of the following valid values can be specified:
O Generate a graph only. Do not format the table contents.
T Include total cost under each operator in the graph.
I Include I/O cost under each operator in the graph.
C Include the expected output cardinality (number of tuples) of each
operator in the graph.
-l Respect case when processing package names.
-n name
Name of the source of the explain request (SOURCE_NAME).
-s schema
Schema or qualifier of the source of the explain request
(SOURCE_SCHEMA).
-o outfile
Output file name.
-t Direct the output to the terminal.
-u userID password
When connecting to a database, use the provided user ID and password.
Both the user ID and password must be valid according to naming
conventions and be recognized by the database.
-w timestamp
Explain time stamp. Specify -1 to obtain the latest explain request.
-# sectnbr
Section number in the source. To request all sections, specify zero.
-h Display help information. When this option is specified, all other options
are ignored, and only the help information is displayed.
Usage notes:
You will be prompted for any parameter values that are not supplied, or that are
incompletely specified, except in the case of the -h and the -l options.
If an explain table schema is not provided, the value of the environment variable
USER is used as the default. If this variable is not found, the user is prompted for
an explain table schema.
Source name, source schema, and explain time stamp can be supplied in LIKE
predicate form, which allows the percent sign (%) and the underscore (_) to be
used as pattern matching characters to select multiple sources with one invocation.
For the latest explained statement, the explain time can be specified as -1.
If -o is specified without a file name, and -t is not specified, the user is prompted
for a file name (the default name is [Link]). If neither -o nor -t is specified,
the user is prompted for a file name (the default option is terminal output). If -o
and -t are both specified, the output is directed to the terminal.
Related concepts:
v “Explain tools” on page 190
v “Guidelines for using explain information” on page 191
v “Guidelines for capturing explain information” on page 198
| For the following examples, computer 1 is called bar and is running AIX. The
| owner of this machine is roecken. The database on bar is called zample. Computer
| 2 is called dps. This machine is also running AIX, and is owned by regress9
| PASSWORDACCESS = generate:
| Computer 1:
| 1. Set up the database for log archiving to TSM. Update the database
| configuration parameter logarchmeth1 for the zample database:
| bar:/home/roecken> db2 update db cfg for zample using LOGARCHMETH1 tsm
| The following information is returned:
| DB20000I The UPDATE DATABASE CONFIGURATION command completed successfully.
| Note: Before updating the database configuration, you may have to take an
| offline backup of the database.
| 2. Take an online backup of the database:
| db2 backup db zample online use tsm
| The following information is returned:
| Backup successful. The timestamp for this backup image is : 20040216151025
| 3. Connect to the zample database, then create a table in it.
| 4. Load data into the new table. In this example, the table is called a, and the data
| is being loaded from a delimited ASCII file called mr. The COPY YES option is
| specified to make a copy of the data that is loaded, and the USE TSM option
| specifies that the copy of the data is stored on Tivoli Storage Manager.
| Note: You can only specify the COPY YES option if the database is enabled for
| rollforward recovery; that is, the logretain or userexit database
| configuration parameter (or both) must be enabled for the database.
| bar:/home/roecken> db2 load from mr of del modified by noheader replace
| into a copy yes use tsm
| The utility returns a series of messages to indicate its progress:
| SQL3109N The utility is beginning to load data from file "/home/roecken/mr".
|
| SQL3500W The utility is beginning the "LOAD" phase at time "02/16/2004
| 15:12:13.392633".
|
| SQL3519W Begin Load Consistency Point. Input record count = "0".
|
| SQL3520W Load Consistency Point was successful.
|
| SQL3110N The utility has completed processing. "1" rows were read from the
| input file.
|
| SQL3519W Begin Load Consistency Point. Input record count = "1".
| Computer 2:
| Computer 2, dps, is not yet set up. A db2adutl query on dps for the zample
| database returns the following results:
| dps:/home/regress9> db2adutl query db zample
| --- Database directory is empty ---
| Warning: There are no file spaces created by DB2 on the ADSM server
| Warning: No DB2 backup images found in ADSM for any alias.
|
|
| dps:/home/regress9> db2adutl query db zample nodename bar owner roecken
| --- Database directory is empty ---
|
| Query for database ZAMPLE
|
|
| Retrieving FULL DATABASE BACKUP information.
| 1 Time: 20040216151025 Oldest log: [Link] DB Partition Number: 0
| Sessions: 1
|
|
| Retrieving INCREMENTAL DATABASE BACKUP information.
| No INCREMENTAL DATABASE BACKUP images found for ZAMPLE
|
|
| Retrieving DELTA DATABASE BACKUP information.
| No DELTA DATABASE BACKUP images found for ZAMPLE
|
|
| Retrieving TABLESPACE BACKUP information.
| No TABLESPACE BACKUP images found for ZAMPLE
|
|
| Retrieving INCREMENTAL TABLESPACE BACKUP information.
| No INCREMENTAL TABLESPACE BACKUP images found for ZAMPLE
|
|
| Retrieving DELTA TABLESPACE BACKUP information.
| No DELTA TABLESPACE BACKUP images found for ZAMPLE
|
|
| Retrieving LOAD COPY information.
| 1 Time: 20040216151213
|
|
| Retrieving LOG ARCHIVE information.
| Log file: [Link], Chain Num: 0, DB Partition Number: 0,
| Taken at: 2004-02-16-15.10.38
| The zample database does not yet exist on the dps computer.
| 1. Restore the zample database to the dps computer:
| dps:/home/regress9> db2 restore db zample use tsm options
| "’-fromnode=bar -fromowner=roecken’" without prompting
| The following information is returned:
| DB20000I The RESTORE DATABASE command completed successfully.
| For db2adutl, update the [Link] (on Windows-based platforms, the [Link])
| and add NODENAME bar (because bar is the name of the source computer) to the
| server clause:
| dps:/home/regress9> db2adutl query db zample nodename bar
| owner roecken password *******
| Related reference:
| v “db2adutl - Managing DB2 objects within TSMCommand” in the Command
| Reference
| v “logarchopt1 - Primary log archive options” on page 401
| v “vendoropt - Vendor options” on page 407
You can access additional DB2 Universal Database™ technical information such as
technotes, white papers, and Redbooks™ online at [Link]®. Access the DB2
Information Management software library site at
[Link]/software/data/pubs/.
| The Information Center is updated more frequently than either the PDF or the
| hardcopy books. To get the most current DB2 technical information, install the
| documentation updates as they become available or go to the DB2 Information
| Center at the [Link] site.
Related concepts:
v “CLI sample programs” in the CLI Guide and Reference, Volume 1
v “Java sample programs” in the Application Development Guide: Building and
Running Applications
v “DB2 Information Center” on page 596
Related tasks:
v “Invoking contextual help from a DB2 tool” on page 613
Related reference:
v “DB2 PDF and printed documentation” on page 607
The DB2 Information Center has the following features if you view it in Mozilla 1.0
or later or Microsoft® Internet Explorer 5.5 or later. Some features require you to
enable support for JavaScript™:
Flexible installation options
You can choose to view the DB2 documentation using the option that best
meets your needs:
v To effortlessly ensure that your documentation is always up to date, you
can access all of your documentation directly from the DB2 Information
Center hosted on the IBM® Web site at
[Link]
v To minimize your update efforts and keep your network traffic within
your intranet, you can install the DB2 documentation on a single server
on your intranet
v To maximize your flexibility and reduce your dependence on network
connections, you can install the DB2 documentation on your own
computer
Search
| You can search all of the topics in the DB2 Information Center by entering
| a search term in the Search text field. You can retrieve exact matches by
| enclosing terms in quotation marks, and you can refine your search with
| wildcard operators (*, ?) and Boolean operators (AND, NOT, OR).
Task-oriented table of contents
| You can locate topics in the DB2 documentation from a single table of
| contents. The table of contents is organized primarily by the kind of tasks
| you may want to perform, but also includes entries for product overviews,
| goals, reference information, an index, and a glossary.
| v Product overviews describe the relationship between the available
| products in the DB2 family, the features offered by each of those
| products, and up to date release information for each of these products.
| v Goal categories such as installing, administering, and developing include
| topics that enable you to quickly complete tasks and develop a deeper
| understanding of the background information for completing those
| tasks.
For iSeries™ technical information, refer to the IBM eServer™ iSeries information
center at [Link]/eserver/iseries/infocenter/.
Related concepts:
v “DB2 Information Center installation scenarios” on page 597
Related tasks:
v “Updating the DB2 Information Center installed on your computer or intranet
server” on page 605
v “Displaying topics in your preferred language in the DB2 Information Center”
on page 606
v “Invoking the DB2 Information Center” on page 604
v “Installing the DB2 Information Center using the DB2 Setup wizard (UNIX)” on
page 600
v “Installing the DB2 Information Center using the DB2 Setup wizard (Windows)”
on page 602
| Tsu-Chen owns a factory in a small town that does not have a local ISP to provide
| him with Internet access. He purchased DB2 Universal Database™ to manage his
| inventory, his product orders, his banking account information, and his business
| expenses. Never having used a DB2 product before, Tsu-Chen needs to learn how
| to do so from the DB2 product documentation.
| After installing DB2 Universal Database on his computer using the typical
| installation option, Tsu-Chen tries to access the DB2 documentation. However, his
| browser gives him an error message that the page he tried to open cannot be
| found. Tsu-Chen checks the installation manual for his DB2 product and discovers
| that he has to install the DB2 Information Center if he wants to access DB2
| documentation on his computer. He finds the DB2 Information Center CD in the
| media pack and installs it.
| From the application launcher for his operating system, Tsu-Chen now has access
| to the DB2 Information Center and can learn how to use his DB2 product to
| increase the success of his business.
| Scenario: Accessing the DB2 Information Center on the IBM Web site:
| Most of the businesses at which Colin teaches have Internet access. This situation
| influenced Colin’s decision to configure his mobile computer to access the DB2
| Information Center on the IBM Web site when he installed the latest version of
| DB2 Universal Database. This configuration allows Colin to have online access to
| the latest DB2 documentation during his seminars.
| Colin enjoys the flexibility of always having a copy of DB2 documentation at his
| disposal. Using the db2set command, he can easily configure the registry variables
| on his mobile computer to access the DB2 Information Center on either the IBM
| Web site, or his mobile computer, depending on his situation.
| Eva works as a senior database administrator for a life insurance company. Her
| administration responsibilities include installing and configuring the latest version
| of DB2 Universal Database on the company’s UNIX® database servers. Her
| company recently informed its employees that, for security reasons, it would not
| provide them with Internet access at work. Because her company has a networked
| environment, Eva decides to install a copy of the DB2 Information Center on an
| intranet server so that all employees in the company who use the company’s data
| warehouse on a regular basis (sales representatives, sales managers, and business
| analysts) have access to DB2 documentation.
| Eva instructs her database team to install the latest version of DB2 Universal
| Database on all of the employee’s computers using a response file, to ensure that
| each computer is configured to access the DB2 Information Center using the host
| name and the port number of the intranet server.
| Related concepts:
| v “DB2 Information Center” on page 596
| Related tasks:
| v “Updating the DB2 Information Center installed on your computer or intranet
| server” on page 605
| v “Installing the DB2 Information Center using the DB2 Setup wizard (UNIX)” on
| page 600
| v “Installing the DB2 Information Center using the DB2 Setup wizard (Windows)”
| on page 602
| v “Setting the location for accessing the DB2 Information Center: Common GUI
| help”
| Related reference:
| v “db2set - DB2 Profile Registry Command” in the Command Reference
| Prerequisites:
| This section lists the hardware, operating system, software, and communication
| requirements for installing the DB2 Information Center on UNIX computers.
| v Hardware requirements
| You require one of the following processors:
| – PowerPC (AIX)
| – HP 9000 (HP-UX)
| – Intel 32–bit (Linux)
| – Solaris UltraSPARC computers (Solaris Operating Environment)
| v Operating system requirements
| You require one of the following operating systems:
| – IBM AIX 5.1 (on PowerPC)
| – HP-UX 11i (on HP 9000)
| – Red Hat Linux 8.0 (on Intel 32–bit)
| – SuSE Linux 8.1 (on Intel 32–bit)
| – Sun Solaris Version 8 (on Solaris Operating Environment UltraSPARC
| computers)
| Note: The DB2 Information Center runs on a subset of the UNIX operating
| systems on which DB2 clients are supported. It is therefore recommended
| that you either access the DB2 Information Center from the IBM Web site,
| or that you install and access the DB2 Information Center on an intranet
| server.
| v Software requirements
| – The following browser is supported:
| - Mozilla Version 1.0 or greater
| v The DB2 Setup wizard is a graphical installer. You must have an implementation
| of the X Window System software capable of rendering a graphical user
| interface for the DB2 Setup wizard to run on your computer. Before you can run
| the DB2 Setup wizard you must ensure that you have properly exported your
| display. For example, enter the following command at the command prompt:
| export DISPLAY=[Link]:0.
| v Communication requirements
| – TCP/IP
| Procedure:
| To install the DB2 Information Center using the DB2 Setup wizard:
| You can also install the DB2 Information Center using a response file.
| The [Link] file captures all DB2 product installation information, including
| errors. The [Link] file records all DB2 product installations on your
| computer. DB2 appends the [Link] file to the [Link] file. The
| [Link] file captures any error output that is returned by Java, for example,
| exceptions and trap information.
| When the installation is complete, the DB2 Information Center will be installed in
| one of the following directories, depending upon your UNIX operating system:
| v AIX: /usr/opt/db2_08_01
| v HP-UX: /opt/IBM/db2/V8.1
| v Linux: /opt/IBM/db2/V8.1
| v Solaris Operating Environment: /opt/IBM/db2/V8.1
| Related concepts:
| v “DB2 Information Center” on page 596
| v “DB2 Information Center installation scenarios” on page 597
Appendix F. DB2 Universal Database technical information 601
| Related tasks:
| v “Installing DB2 using a response file (UNIX)” in the Installation and Configuration
| Supplement
| v “Updating the DB2 Information Center installed on your computer or intranet
| server” on page 605
| v “Displaying topics in your preferred language in the DB2 Information Center”
| on page 606
| v “Invoking the DB2 Information Center” on page 604
| v “Installing the DB2 Information Center using the DB2 Setup wizard (Windows)”
| on page 602
| Installing the DB2 Information Center using the DB2 Setup wizard
| (Windows)
| DB2 product documentation can be accessed in three ways: on the IBM Web site,
| on an intranet server, or on a version installed on your computer. By default, DB2
| products access DB2 documentation on the IBM Web site. If you want to access the
| DB2 documentation on an intranet server or on your own computer, you must
| install the DB2 documentation from the DB2 Information Center CD. Using the DB2
| Setup wizard, you can define your installation preferences and install the DB2
| Information Center on a computer that uses a Windows operating system.
| Prerequisites:
| This section lists the hardware, operating system, software, and communication
| requirements for installing the DB2 Information Center on Windows.
| v Hardware requirements
| You require one of the following processors:
| – 32-bit computers: a Pentium or Pentium compatible CPU
| v Operating system requirements
| You require one of the following operating systems:
| – Windows 2000
| – Windows XP
| Note: The DB2 Information Center runs on a subset of the Windows operating
| systems on which DB2 clients are supported. It is therefore recommended
| that you either access the DB2 Information Center on the IBM Web site, or
| that you install and access the DB2 Information Center on an intranet
| server.
| v Software requirements
| – The following browsers are supported:
| - Mozilla 1.0 or greater
| - Internet Explorer Version 5.5 or 6.0 (Version 6.0 for Windows XP)
| v Communication requirements
| – TCP/IP
| Restrictions:
| v You require an account with administrative privileges to install the DB2
| Information Center.
| To install the DB2 Information Center using the DB2 Setup wizard:
| 1. Log on to the system with the account that you have defined for the DB2
| Information Center installation.
| 2. Insert the CD into the drive. If enabled, the auto-run feature starts the IBM
| DB2 Setup Launchpad.
| 3. The DB2 Setup wizard determines the system language and launches the
| setup program for that language. If you want to run the setup program in a
| language other than English, or the setup program fails to auto-start, you can
| start the DB2 Setup wizard manually.
| To start the DB2 Setup wizard manually:
| a. Click Start and select Run.
| b. In the Open field, type the following command:
| x:\[Link] /i 2-letter language identifier
| You can install the DB2 Information Center using a response file. You can also use
| the db2rspgn command to generate a response file based on an existing
| installation.
| For information on errors encountered during installation, see the [Link] and
| [Link] files located in the ’My Documents’\DB2LOG\ directory. The location of the
| ’My Documents’ directory will depend on the settings on your computer.
| The [Link] file captures the most recent DB2 installation information. The
| [Link] captures the history of DB2 product installations.
| Related tasks:
| v “Installing a DB2 product using a response file (Windows)” in the Installation and
| Configuration Supplement
| v “Updating the DB2 Information Center installed on your computer or intranet
| server” on page 605
| v “Displaying topics in your preferred language in the DB2 Information Center”
| on page 606
| v “Invoking the DB2 Information Center” on page 604
| v “Installing the DB2 Information Center using the DB2 Setup wizard (UNIX)” on
| page 600
| Related reference:
| v “db2rspgn - Response File Generator Command (Windows)” in the Command
| Reference
You can invoke the DB2 Information Center from one of the following places:
v Computers on which a DB2 UDB client or server is installed
v An intranet server or local computer on which the DB2 Information Center
installed
v The IBM Web site
Prerequisites:
Procedure:
To invoke the DB2 Information Center on a computer on which a DB2 UDB client
or server is installed:
v From the Start Menu (Windows operating system): Click Start — Programs —
IBM DB2 — Information — Information Center.
v From the command line prompt:
– For Linux and UNIX operating systems, issue the db2icdocs command.
– For the Windows operating system, issue the [Link] command.
To open the DB2 Information Center on the IBM Web site in a Web browser:
v Open the Web page at [Link]/infocenter/db2help/.
Related concepts:
v “DB2 Information Center” on page 596
v “DB2 Information Center installation scenarios” on page 597
Related tasks:
v “Displaying topics in your preferred language in the DB2 Information Center”
on page 606
v “Invoking contextual help from a DB2 tool” on page 613
v “Updating the DB2 Information Center installed on your computer or intranet
server” on page 605
v “Invoking command help from the command line processor” on page 615
v “Setting the location for accessing the DB2 Information Center: Common GUI
help”
Related reference:
v “HELP Command” in the Command Reference
Prerequisites:
Procedure:
Related concepts:
v “DB2 Information Center installation scenarios” on page 597
Related tasks:
v “Invoking the DB2 Information Center” on page 604
v “Installing the DB2 Information Center using the DB2 Setup wizard (UNIX)” on
page 600
v “Installing the DB2 Information Center using the DB2 Setup wizard (Windows)”
on page 602
| Procedure:
| Note: Adding a language does not guarantee that the computer has the fonts
| required to display the topics in the preferred language.
| v To move a language to the top of the list, select the language and click the
| Move Up button until the language is first in the list of languages.
| 3. Refresh the page to display the DB2 Information Center in your preferred
| language.
The following tables describe, for each book in the DB2 library, the information
needed to order the hard copy, or to print or view the PDF for that book. A full
description of each of the books in the DB2 library is available from the IBM
Publications Center at [Link]/shop/publications/order
| Administration information
The information in these books covers those topics required to effectively design,
implement, and maintain DB2 databases, data warehouses, and federated systems.
Tutorial information
Tutorial information introduces DB2 features and teaches how to perform various
tasks.
Table 79. Tutorial information
Name Form number PDF file name
Business Intelligence Tutorial: No form number db2tux81
Introduction to the Data
Warehouse
Business Intelligence Tutorial: No form number db2tax81
Extended Lessons in Data
Warehousing
Information Catalog Center No form number db2aix81
Tutorial
Video Central for e-business No form number db2twx81
Tutorial
Visual Explain Tutorial No form number db2tvx81
Release notes
The release notes provide additional information specific to your product’s release
and FixPak level. The release notes also provide summaries of the documentation
updates incorporated in each release, update, and FixPak.
Table 81. Release notes
Name Form number PDF file name
DB2 Release Notes See note. See note.
DB2 Installation Notes Available on product Not available.
CD-ROM only.
To view the Release Notes in text format on UNIX-based platforms, see the
[Link] file. This file is located in the DB2DIR/Readme/%L directory,
where %L represents the locale name and DB2DIR represents:
v For AIX operating systems: /usr/opt/db2_08_01
v For all other UNIX-based operating systems: /opt/IBM/db2/V8.1
Related concepts:
v “DB2 documentation and help” on page 595
Related tasks:
v “Printing DB2 books from PDF files” on page 612
v “Ordering printed DB2 books” on page 612
v “Invoking contextual help from a DB2 tool” on page 613
Prerequisites:
Ensure that you have Adobe Acrobat Reader installed. If you need to install Adobe
Acrobat Reader, it is available from the Adobe Web site at [Link]
Procedure:
Related concepts:
v “DB2 Information Center” on page 596
Related tasks:
v “Mounting the CD-ROM (AIX)” in the Quick Beginnings for DB2 Servers
v “Mounting the CD-ROM (HP-UX)” in the Quick Beginnings for DB2 Servers
v “Mounting the CD-ROM (Linux)” in the Quick Beginnings for DB2 Servers
v “Ordering printed DB2 books” on page 612
v “Mounting the CD-ROM (Solaris Operating Environment)” in the Quick
Beginnings for DB2 Servers
Related reference:
v “DB2 PDF and printed documentation” on page 607
Procedure:
| Printed books can be ordered in some countries or regions. Check the IBM
| Publications website for your country or region to see if this service is available in
| your country or region. When the publications are available for ordering, you can:
| v Contact your IBM authorized dealer or marketing representative. To find a local
| IBM representative, check the IBM Worldwide Directory of Contacts at
| [Link]/planetwide
| v Phone 1-800-879-2755 in the United States or 1-800-IBM-4YOU in Canada.
At the time the DB2 product becomes available, the printed books are the same as
those that are available in PDF format on the DB2 PDF Documentation CD. Content
in the printed books that appears in the DB2 Information Center CD is also the
same. However, there is some additional content available in DB2 Information
Center CD that does not appear anywhere in the PDF books (for example, SQL
Administration routines and HTML samples). Not all books available on the DB2
PDF Documentation CD are available for ordering in hardcopy.
Note: The DB2 Information Center is updated more frequently than either the PDF
or the hardcopy books; install documentation updates as they become
available or refer to the DB2 Information Center at
[Link] to get the most current
information.
Related tasks:
v “Printing DB2 books from PDF files” on page 612
Related reference:
v “DB2 PDF and printed documentation” on page 607
Procedure:
Related tasks:
v “Invoking the DB2 Information Center” on page 604
v “Invoking message help from the command line processor” on page 614
v “Invoking command help from the command line processor” on page 615
v “Invoking SQL state help from the command line processor” on page 615
v “Access to the DB2 Information Center: Concepts help”
v “How to use the DB2 UDB help: Common GUI help”
v “Setting the location for accessing the DB2 Information Center: Common GUI
help”
v “Setting up access to DB2 contextual help and documentation: Common GUI
help”
| Procedure:
| To invoke message help, open the command line processor and enter:
| ? XXXnnnnn
| Related concepts:
| v “Introduction to messages” in the Message Reference Volume 1
| Related reference:
| v “db2 - Command Line Processor Invocation Command” in the Command
| Reference
| Procedure:
| To invoke command help, open the command line processor and enter:
| ? command
| For example, ? catalog displays help for all of the CATALOG commands, while ?
| catalog database displays help only for the CATALOG DATABASE command.
| Related tasks:
| v “Invoking contextual help from a DB2 tool” on page 613
| v “Invoking the DB2 Information Center” on page 604
| v “Invoking message help from the command line processor” on page 614
| v “Invoking SQL state help from the command line processor” on page 615
| Related reference:
| v “db2 - Command Line Processor Invocation Command” in the Command
| Reference
| Procedure:
| To invoke SQL state help, open the command line processor and enter:
| ? sqlstate or ? class code
| where sqlstate represents a valid five-digit SQL state and class code represents the
| first two digits of the SQL state.
| For example, ? 08003 displays help for the 08003 SQL state, and ? 08 displays help
| for the 08 class code.
| Related tasks:
| v “Invoking the DB2 Information Center” on page 604
| v “Invoking message help from the command line processor” on page 614
| v “Invoking command help from the command line processor” on page 615
DB2 tutorials
The DB2® tutorials help you learn about various aspects of DB2 Universal
Database. The tutorials provide lessons with step-by-step instructions in the areas
of developing applications, tuning SQL query performance, working with data
warehouses, managing metadata, and developing Web services using DB2.
You can view the XHTML versions of the tutorials from the Information Center at
[Link]
Some tutorial lessons use sample data or code. See each tutorial for a description
of any prerequisites for its specific tasks.
Related concepts:
v “DB2 Information Center” on page 596
v “Introduction to problem determination - DB2 Technical Support tutorial” in the
Troubleshooting Guide
Accessibility
Accessibility features help users with physical disabilities, such as restricted
mobility or limited vision, to use software products successfully. The following list
specifies the major accessibility features in DB2® Version 8 products:
v All DB2 functionality is available using the keyboard for navigation instead of
the mouse. For more information, see “Keyboard input and navigation.”
v You can customize the size and color of the fonts on DB2 user interfaces. For
more information, see “Accessible display.”
v DB2 products support accessibility applications that use the Java™ Accessibility
API. For more information, see “Compatibility with assistive technologies” on
page 618.
v DB2 documentation is provided in an accessible format. For more information,
see “Accessible documentation” on page 618.
| For more information about using keys or key combinations to perform operations,
| see Keyboard shortcuts and accelerators: Common GUI help.
Keyboard navigation
You can navigate the DB2 tools user interface using keys or key combinations.
For more information about using keys or key combinations to navigate the DB2
Tools, see Keyboard shortcuts and accelerators: Common GUI help.
Keyboard focus
In UNIX® operating systems, the area of the active window where your keystrokes
will have an effect is highlighted.
Accessible display
The DB2 tools have features that improve accessibility for users with low vision or
other visual impairments. These accessibility enhancements include support for
customizable font properties.
For more information about specifying font settings, see Changing the fonts for
menus and text: Common GUI help.
Non-dependence on color
You do not need to distinguish between colors in order to use any of the functions
in this product.
Accessible documentation
Documentation for DB2 is provided in XHTML 1.0 format, which is viewable in
most Web browsers. XHTML allows you to view documentation according to the
display preferences set in your browser. It also allows you to use screen readers
and other assistive technologies.
Syntax diagrams are provided in dotted decimal format. This format is available
only if you are accessing the online documentation using a screen-reader.
Related concepts:
v “Dotted decimal syntax diagrams” on page 618
Related tasks:
v “Keyboard shortcuts and accelerators: Common GUI help”
v “Changing the fonts for menus and text: Common GUI help”
| In dotted decimal format, each syntax element is written on a separate line. If two
| or more syntax elements are always present together (or always absent together),
| they can appear on the same line, because they can be considered as a single
| compound syntax element.
| Each line starts with a dotted decimal number; for example, 3 or 3.1 or 3.1.1. To
| hear these numbers correctly, make sure that your screen reader is set to read out
| punctuation. All the syntax elements that have the same dotted decimal number
| (for example, all the syntax elements that have the number 3.1) are mutually
| exclusive alternatives. If you hear the lines 3.1 USERID and 3.1 SYSTEMID, you
| know that your syntax can include either USERID or SYSTEMID, but not both.
| The dotted decimal numbering level denotes the level of nesting. For example, if a
| syntax element with dotted decimal number 3 is followed by a series of syntax
| elements with dotted decimal number 3.1, all the syntax elements numbered 3.1
| are subordinate to the syntax element numbered 3.
| The following words and symbols are used next to the dotted decimal numbers:
| v ? means an optional syntax element. A dotted decimal number followed by the ?
| symbol indicates that all the syntax elements with a corresponding dotted
| decimal number, and any subordinate syntax elements, are optional. If there is
| only one syntax element with a dotted decimal number, the ? symbol is
| displayed on the same line as the syntax element, (for example 5? NOTIFY). If
| there is more than one syntax element with a dotted decimal number, the ?
| symbol is displayed on a line by itself, followed by the syntax elements that are
| optional. For example, if you hear the lines 5 ?, 5 NOTIFY, and 5 UPDATE, you
| know that syntax elements NOTIFY and UPDATE are optional; that is, you can
| choose one or none of them. The ? symbol is equivalent to a bypass line in a
| railroad diagram.
| v ! means a default syntax element. A dotted decimal number followed by the !
| symbol and a syntax element indicates that the syntax element is the default
| option for all syntax elements that share the same dotted decimal number. Only
| one of the syntax elements that share the same dotted decimal number can
| specify a ! symbol. For example, if you hear the lines 2? FILE, 2.1! (KEEP), and
| 2.1 (DELETE), you know that (KEEP) is the default option for the FILE keyword.
| In this example, if you include the FILE keyword but do not specify an option,
| default option KEEP will be applied. A default option also applies to the next
| higher dotted decimal number. In this example, if the FILE keyword is omitted,
| default FILE(KEEP) is used. However, if you hear the lines 2? FILE, 2.1, 2.1.1!
| (KEEP), and 2.1.1 (DELETE), the default option KEEP only applies to the next
| higher dotted decimal number, 2.1 (which does not have an associated
| keyword), and does not apply to 2? FILE. Nothing is used if the keyword FILE
| is omitted.
| v * means a syntax element that can be repeated 0 or more times. A dotted
| decimal number followed by the * symbol indicates that this syntax element can
| be used zero or more times; that is, it is optional and can be repeated. For
| example, if you hear the line 5.1* data area, you know that you can include one
| Related concepts:
| v “Accessibility” on page 617
| Related tasks:
| v “Keyboard shortcuts and accelerators: Common GUI help”
| Related reference:
| v “How to read the syntax diagrams” in the SQL Reference, Volume 2
IBM may have patents or pending patent applications covering subject matter
described in this document. The furnishing of this document does not give you
any license to these patents. You can send license inquiries, in writing, to:
IBM Director of Licensing
IBM Corporation
North Castle Drive
Armonk, NY 10504-1785
U.S.A.
For license inquiries regarding double-byte (DBCS) information, contact the IBM
Intellectual Property Department in your country/region or send inquiries, in
writing, to:
IBM World Trade Asia Corporation
Licensing
2-31 Roppongi 3-chome, Minato-ku
Tokyo 106, Japan
The following paragraph does not apply to the United Kingdom or any other
country/region where such provisions are inconsistent with local law:
INTERNATIONAL BUSINESS MACHINES CORPORATION PROVIDES THIS
PUBLICATION “AS IS” WITHOUT WARRANTY OF ANY KIND, EITHER
EXPRESS OR IMPLIED, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
WARRANTIES OF NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS
FOR A PARTICULAR PURPOSE. Some states do not allow disclaimer of express or
implied warranties in certain transactions; therefore, this statement may not apply
to you.
Any references in this information to non-IBM Web sites are provided for
convenience only and do not in any manner serve as an endorsement of those Web
sites. The materials at those Web sites are not part of the materials for this IBM
product, and use of those Web sites is at your own risk.
IBM may use or distribute any of the information you supply in any way it
believes appropriate without incurring any obligation to you.
The licensed program described in this document and all licensed material
available for it are provided by IBM under terms of the IBM Customer Agreement,
IBM International Program License Agreement, or any equivalent agreement
between us.
All statements regarding IBM’s future direction or intent are subject to change or
withdrawal without notice, and represent goals and objectives only.
This information may contain examples of data and reports used in daily business
operations. To illustrate them as completely as possible, the examples include the
names of individuals, companies, brands, and products. All of these names are
fictitious, and any similarity to the names and addresses used by an actual
business enterprise is entirely coincidental.
COPYRIGHT LICENSE:
Each copy or any portion of these sample programs or any derivative work must
include a copyright notice as follows:
Trademarks
The following terms are trademarks of International Business Machines
Corporation in the United States, other countries, or both, and have been used in
at least one of the documents in the DB2 UDB documentation library.
ACF/VTAM iSeries
AISPO LAN Distance
AIX MVS
AIXwindows MVS/ESA
AnyNet MVS/XA
APPN [Link]
AS/400 NetView
BookManager OS/390
C Set++ OS/400
C/370 PowerPC
CICS pSeries
Database 2 QBIC
DataHub QMF
DataJoiner RACF
DataPropagator RISC System/6000
DataRefresher RS/6000
DB2 S/370
DB2 Connect SP
DB2 Extenders SQL/400
DB2 OLAP Server SQL/DS
DB2 Information Integrator System/370
DB2 Query Patroller System/390
DB2 Universal Database SystemView
Distributed Relational Tivoli
Database Architecture VisualAge
DRDA VM/ESA
eServer VSE/ESA
Extended Services VTAM
FFST WebExplorer
First Failure Support Technology WebSphere
IBM WIN-OS/2
IMS z/OS
IMS/ESA zSeries
Microsoft, Windows, Windows NT, and the Windows logo are trademarks of
Microsoft Corporation in the United States, other countries, or both.
Intel and Pentium are trademarks of Intel Corporation in the United States, other
countries, or both.
Java and all Java-based trademarks are trademarks of Sun Microsystems, Inc. in the
United States, other countries, or both.
Index 627
database system monitor DB2_LIKE_VARCHAR 502 db2expln tool (continued)
default database system monitor DB2_MINIMIZE_LISTPREFETCH 502 information displayed
switches configuration DB2_MMAP_READ 506 aggregation 570
parameter 455 DB2_MMAP_WRITE 506 data stream 568
database territory code configuration DB2_NEW_CORR_SQ_FF 502 insert, update, delete 568
parameter 424 DB2_NEWLOGPATH2 518 join 566
database_consistent configuration DB2_NO_FORK_CHECK 506 miscellaneous 574
parameter 429 DB2_NO_MPFA_FOR_NEW_DB 506 parallel processing 571
database_level configuration DB2_NUM_FAILOVER_NODES 500 table access 559
parameter 424 DB2_OBJECT_TABLE_ENTRIES 506 temporary table 564
database_memory configuration DB2_OVERRIDE_BPF 506 output 558
parameter 338 DB2_PARALLEL_IO 492 output samples
database-managed space (DMS) DB2_PARTITIONEDLOAD__DEFAULT 500 description 576
description 15 DB2_PINNED_BP 506 for federated database plan 584
table-space address map 17 DB2_PRED_FACTORIZE 502 multipartition plan with full
databases DB2_REDUCED_OPTIMIZATION 502 parallelism 582
autorestart configuration DB2_SCATTERED_IO 506 multipartition plan with
parameter 408 DB2_SELECTIVITY 502 inter-partition parallelism 579
backup_pending configuration DB2_SMS_TRUNC_TMP_TABLE_THRESH 506 no parallelism 576
parameter 429 DB2_SORT_AFTER_TQ 506 single-partition plan with
codepage configuration DB2_TRUSTED_BINDIN 506 intra-partition parallelism 578
parameter 423 DB2_USE_ALTERNATE_PAGE_CLEANING 506syntax and parameters 552
codeset configuration parameter 424 usage 224 usage notes 557
collating information 424 DB2_USE_PAGE_CONTAINER DB2GRAPHICUNICODESERVER 490
configuration parameter _TAG 492 DB2INCLUDE 490
summary 323 DB2_USE_PAGE_CONTAINER_TAG 492 DB2INSTANCE 492
maximum number of concurrently DB2_VENDOR_INI 518 DB2INSTDEF 490
active databases 460 DB2_VI_DEVICE 496 DB2INSTOWNER 490
release level configuration DB2_VI_ENABLE 496 DB2INSTPROF 492
parameter 425 DB2_VI_VIPL 496 DB2IQTIME 499
territory code configuration DB2_VIEW_REOPT_VALUES 490 DB2JD_PORT_NUMBER 496
parameter 424 DB2_XBSA_LIBRARY 518 DB2LDAP_BASEDN 518
territory configuration parameter 425 DB2ACCOUNT 490 DB2LDAP_CLIENT_PROVIDER 518
DATALINK data type DB2ADMINSERVER 518 DB2LDAP_SEARCH_SCOPE 518
configuration parameter 426 db2adutl command DB2LDAPCACHE 518
DB2 architecture overview 9 cross-node recovery example 589 DB2LDAPHOST 518
DB2 books db2advis 7, 246 DB2LIBPATH 492
printing PDF files 612 DB2AFFINITIES 506 DB2LOADREC 518
DB2 Information Center 596 DB2ASSUMEUPDATE 506 DB2LOCALE 490
invoking 604 DB2ATLD_PWFILE 500 DB2LOCK_TO_RB 518
DB2 tutorials 615 db2batch benchmarking tool DB2MAXFSCRSEARCH 506
DB2_ALLOCATION_SIZE 506 creating tests 305 DB2MEMDISCLAIM 506
DB2_ANTIJOIN 502 examples 307 DB2MEMMAXFREE 506
DB2_APM_PERFORMANCE 506 DB2BIDI 490 DB2NBADAPTERS 496
DB2_AVOID_PREFETCH 506 DB2BPVARS 506 DB2NBCHECKUPTIME 496
DB2_AWE 506 DB2BQTIME 499 DB2NBDISCOVERRCVBUFS 490
DB2_BINSORT 506 DB2BQTRY 499 DB2NBINTRLISTENS 496
DB2_CLPPROMPT 499 DB2CHECKCLIENTINTERVAL 496 DB2NBRECVBUFFSIZE 496
DB2_CORRELATED_PREDICATES 502 DB2CHGPWD_ESE 500 DB2NBRECVNCBS 496
DB2_DJ_COMM 518 DB2CHKPTR 506 DB2NBRESOURCES 496
DB2_DOCHOST 518 DB2CHKSQLDA 506 DB2NBSENDNCBS 496
DB2_DOCPORT 518 DB2CLIINIPATH 518 DB2NBSESSIONS 496
DB2_ENABLE_BUFPD 506 DB2CODEPAGE 490 DB2NBXTRANCBS 496
DB2_ENABLE_LDAP 518 DB2COMM 496 DB2NODE 492
DB2_EVALUNCOMMITTED 506 DB2CONNECT_IN_APP_PROCESS 492 exported when adding server 283,
DB2_EXTENDED_OPTIMIZATION 506 DB2DBDFT 490 284, 285
DB2_FALLBACK 518 DB2DBMSADDR 490 DB2NOEXITLIST 518
DB2_FMP_COMM_HEAPSZ 518 DB2DEFPREP 518 DB2NTMEMSIZE 506
DB2_FORCE_FCM_BP 500 DB2DISCOVERYTIME 490 DB2NTNOCACHE 506
DB2_FORCE_NLS_CACHE 496 DB2DMNBCKCTLR 518 DB2NTPRICLASS 506
DB2_GRP_LOOKUP 518 DB2DOMAINLIST 492 DB2NTWORKSET 506
DB2_HASH_JOIN 502 db2empfa command 14 DB2PATH 492
DB2_INDEX_TYPE2 490 DB2ENVLIST 492 DB2PORTRANGE 500
DB2_INLIST_TO_NLJN 502 db2exfmt tool 587 DB2PRIORITIES 506
DB2_KEEPTABLELOCK 506 db2expln tool DB2REMOTEPREG 518
DB2_LGPAGE_BP 506 block and RID preparation DB2RETRY 496
DB2_LIC_STAT_SIZE 490 information 569 DB2RETRYTIME 496
Index 629
fcm_num_anchors configuration hadr_remote_svc configuration initial number of fenced processes
parameter 444 parameter 412 configuration parameter 389
fcm_num_buffers configuration hadr_syncmode configuration inserting data
parameter 444 parameter 412 process for 26
fcm_num_connect configuration hadr_timeout configuration when table clustered on index 26
parameter 446 parameter 413 installing
fcm_num_rqb configuration hash join Information Center 597, 600, 602
parameter 446 described 157 instance memory configuration
fed_noauth configuration parameter 468 tuning performance of 157 parameter 364
federated configuration parameter 458 health monitoring configuration instance_memory configuration
federated databases parameter 453 parameter 364
analyzing where queries health_mon configuration parameter 453 intra_parallel configuration
evaluated 182 help parameter 449
compiler phases 178 displaying 604, 606 intra-partition parallelism
concurrency control for 39 for commands optimization strategies for 173
db2expln output for query in 584 invoking 615 invoking
global analysis of queries on 186 for messages command help 615
global optimization in 184 invoking 614 message help 614
pushdown analysis 178 for SQL statements SQL statement help 615
query information 573 invoking 615 IS (intent share) mode 47
server options 94 HTML documentation isolation levels
system support configuration updating 605 effect on performance 40
parameter 458 locks for concurrency control 46
fenced_pool configuration specifying 43
parameter 386
first active log file configuration
I statement-level 43
I/O parallelism
parameter 391
managing 235
FOR FETCH ONLY clause
in query tuning 77
INCLUDE clause J
effect on space required for Java Development Kit installation path
FOR READ ONLY clause
indexes 18 (DAS) configuration parameter 482
in query tuning 77
index re-creation time configuration Java Development Kit installation path
free space control record (FSCR)
parameter 413 configuration parameter 459
in MDC tables 21
index scans java_heap_sz configuration
in standard tables 18
accessing data through 148 parameter 365
previous leaf pointers 23 jdk_64_path configuration
search processes 23 parameter 482
G usage 23 jdk_path configuration parameter 459
governor tool indexes jdk_path DAS configuration
configuration file example 275 advantages of 244 parameter 482
configuration file rule block index-scan lock mode 65 joins
descriptions 269 cluster ratio 153 broadcast inner-table 165
configuring 268 clustering 18 broadcast outer-table 165
daemon described 267 collecting catalog statistics on 100 collocated 165
described 265 data-access methods using 151 db2expln information displayed
log files created by 276 defragmentation, online 254 for 566
queries against log files 280 detailed statistics data collected 120 described 156
rule elements 271 effect of type on next-key locking 70 eliminating redundancy 140
starting and stopping 266 index re-creation time configuration hash, described 157
group_plugin configuration parameter 413 in partitioned databases 165
parameter 468 managing 244, 251 merge, described 157
groupheap_ratio configuration managing for MDC tables 21 methods, listed 157
parameter 348 managing for standard tables 18 nested-loop, described 157
grouping effect on access plan 171 performance tips for 248 optimizer strategies for optimal 160
planning 246 shared aggregation 140
reorganizing 252 subquery transformation by
H rules for updating statistics
manually 131
optimizer 140
table-queue strategy in partitioned
hadr_db_role configuration
scans 23 databases 164
parameter 409
structure 23 types
hadr_local_host configuration
type-2 described 251 directed inner-table 165
parameter 410
when to create 246 directed outer-table 165
hadr_local_svc configuration
wizards to help design 201
parameter 410
indexrec configuration parameter 413
hadr_remote_host configuration
parameter 411
Information Center
installing 597, 600, 602
K
hadr_remote_inst configuration keepfenced configuration parameter 388
initial number of agents in pool
parameter 411
configuration parameter 385
Index 631
memory (continued) nodetype configuration parameter 459 overhead
statement heap size configuration notifylevel configuration parameter 453 row blocking to reduce 80
parameter 357 num_db_backups configuration
tuning parameters that affect 218 parameter 415
when allocated 211
memory model
num_estore_segs configuration parameter
description 373
P
page cleaners
database-manager shared for memory management 32
tuning number of 223
memory 213 num_freqvalues configuration
pages, data 18
described 32 parameter 434
parallel processing, information displayed
memory requirements num_initfenced configuration
by db2expln output 571
FCM buffer pool 215 parameter 389
parallelism
merge join 157 num_iocleaners configuration
effect of
message help parameter 374
dft_degree configuration
invoking 614 num_ioservers configuration
parameter 88
methods parameter 375
intra_parallel configuration
nested-loop join 157 num_poolagents configuration
parameter 88
min_dec_div_3 configuration parameter 385
max_querydegree configuration
parameter 359 num_quantiles configuration
parameter 88
min_priv_mem configuration parameter 435
enable intra-partition parallelism
parameter 351 numarchretry configuration
configuration parameter 449
mincommit configuration parameter 403 parameter 405
I/O
MINPCTUSED clause number of commits to group
managing 235
for online index defragmentation 18 configuration parameter 403
server configuration for 233
mirror log path configuration number of database backups
intra-partition
parameter 395 configuration parameter 415
optimization strategies 173
mirrorlogpath configuration numdb configuration parameter 460
maximum query degree of parallelism
parameter 395 effect on memory use 211
configuration parameter 450
modeling application performance for memory management 32
non-SMP environments 88
using catalog statistics 125 numinitagents configuration
setting degree of 88
using manually adjusted catalog parameter 385
partition groups, effect on query
statistics 124 numsegs configuration parameter 376
optimization 91
mon_heap_sz configuration NW (next key weak exclusive) mode 47
partitioned databases
parameter 366
data redistribution, error
monitor switches
recovery 294
updating 262
monitoring
O decorrelation of a query 143
online errors when adding nodes 287
how to 262
help, accessing 613 join methods in 165
multidimensional clustering (MDC)
operations join strategies in 164
management of tables and
merged or moved by optimizer 139 replicated materialized query tables
indexes 21
optimization in 162
optimization strategies for 175
intra-partition parallelism 173 partitions
multipage_alloc configuration
strategies for MDC tables 175 adding
parameter 429
optimization classes to a running system 283
effect on memory 14
choosing 72 to a stopped system 285
setting for SMS table spaces 14
listed and described 73 to NT system 284
multisite update 39
setting 76 dropping 288
OPTIMIZE FOR clause pckcachesz configuration parameter 343
in query tuning 77 PCTFREE clause
N optimizer to retain space for clustering 18
nested-loop join 157 access plan performance
NetBIOS effect of sorting and adjusting optimization class 76
workstation name configuration grouping 171 db2batch benchmarking tool 305
parameter 439 for column correlation 145 developing improvement process 5
newlogpath configuration index access methods 151 disk-storage factors 12
parameter 396 using index 148 elements of 3
next-key locks distribution statistics, use of 115 federated database systems 178
converting index to minimize 252 joins limits to tuning 6
index type, effects 70 described 156 query optimization using the REOPT
type-2 indexes 251 in partitioned database 165 bind option 147
nname configuration parameter 439 strategies for optimal 160 tuning 3
node 39 query rewriting methods 139 quick-start tips 7
node connection retries configuration ordering DB2 books 612 user input for 6
parameter 447 overflow records point-in-time monitoring 262
nodes in standard tables 18 pool size for agents, controlling 385
connection elapse time 443 performance effect 240 precompiling
coordinating agents, maximum 379 overflowlogpath configuration isolation level 43
maximum time difference among 448 parameter 398
Index 633
registry variables (continued) RUNSTATS srvcon_auth configuration
DB2NTPRICLASS 506 automatic statistics collection 104, parameter 469
DB2NTWORKSET 506 105 srvcon_gssplugin_list configuration
DB2OPTIONS 490 sampling statistics 101 parameter 470
DB2PATH 492 statistics collected 95 srvcon_pw_plugin configuration
DB2PORTRANCE 500 using 98 parameter 471
DB2PRIORITIES 506 start and stop timeout configuration
DB2REMOTEPREG 518 parameter 448
DB2RETRY 496
DB2RETRYTIME 496
S start_stop_time configuration
parameter 448
SARGable
DB2ROUTINE_DEBUG 518 stat_heap_sz configuration
defined 154
DB2RQTIME 499 parameter 356
sched_enable configuration
DB2SERVICETPINSTANCE 496 statement heap size configuration
parameter 483
DB2SLOGON 490 parameter 357
sched_userid configuration
DB2SORCVBUF 518 statement-level isolation, specifying 43
parameter 484
DB2SORT 518 static SQL
SELECT statement
DB2SOSNDBUF 496 setting optimization class 76
eliminating DISTINCT clauses 143
DB2SYSPLEX_SERVER 496 statistics
prioritizing output for 77
DB2SYSTEM 518 automatic collection 104, 105
seqdetect configuration parameter 376
DB2TCPCONNMGRS 496 Statistics collection
sequential prefetching
DB2TERRITORY 490 sampling 101
described 230
DB2TIMEOUT 490 statistics profile
SET CURRENT QUERY OPTIMIZATION
DB2TRACEFLUSH 490 generating 102
statement 76
DB2TRACENAME 490 stmtheap configuration parameter 357
shadow paging, long objects 25
DB2TRACEON 490 stmtheap configuration parameter, effect
sheapthres configuration parameter 354
DB2TRCSYSERR 490 on query optimization 136
sheapthres_shr configuration
DB2YIELD 490 stored procedures
parameter 344
DLFM_ASNCOPYD_PORT 516 how used 87
SIX (share with intent exclusive)
DLFM_BACKUP_DIR_NAME 516 subqueries
mode 47
DLFM_BACKUP_TARGET_LIBRARY 516 correlated
smtp_server configuration
DLFM_GC_MODE 516 how rewritten 143
parameter 484
DLFM_INSTALL_PATH 516 summary tables
snapshots
DLFM_PORT 516 See materialized query tables. 176
point-in-time monitoring 262
DLFM_START_ASNCOPYD 516 svcename configuration parameter 439
softmax configuration parameter 405
DLFM_TSM_MGMTCLASS 516 sysadm_group configuration
sortheap configuration parameter
overview 489 parameter 472
description 355
release configuration parameter 425 sysctrl_group configuration
effect on query optimization 136
remote data services node name parameter 473
sorting
configuration parameter 439 sysmaint_group configuration
effect on access plan 171
REORG INDEXES command 252 parameter 473
managing 236
REORG TABLE command sysmon_group configuration
sort heap size configuration
choosing reorg method 242 parameter 474
parameter 355
classic, in off-line mode 242 system managed space (SMS)
sort heap threshold configuration
in-place, in on-line mode 242 described 14
parameter 354
REORGANIZE TABLE command
sort heap threshold for shared
indexes and tables 252
sorts 344
reorganizing
tables
space map pages (SMP), DMS table T
spaces 15 table spaces
determining when to 240
spm_log_file_sz configuration DMS 15
restore_pending configuration
parameter 419 effect on query optimization 91
parameter 430
spm_log_path configuration lock types 47
resync_interval configuration
parameter 420 overhead 91
parameter 418
spm_max_resync configuration TRANSFERRATE, setting 91
REXX language
parameter 420 tables
isolation level, specifying 43
spm_name configuration parameter 421 access
roll-forward recovery
SQL compiler information displayed by
definition 25
process description 133 db2expln 559
rollforward utility
SQL Explain 189 paths 60
roll forward pending indicator 430
SQL statement help lock modes
rollfwd_pending configuration
invoking 615 for RID and table scans of MDC
parameter 430
SQL statements tables 62
row blocking
benchmarking 304 for standard tables 60
specifying 80
statement heap size configuration lock types 47
rows
parameter 357 multidimensional clustering 21
lock types 47
SQLDBCON configuration file 315 queues, for join strategies in
rqrioblk configuration parameter 360
srv_plugin_mode configuration partitioned databases 164
parameter 471
Index 635
636 Administration Guide: Performance
Contacting IBM
In the United States, call one of the following numbers to contact IBM:
v 1-800-IBM-SERV (1-800-426-7378) for customer service
v 1-888-426-4343 to learn about available service options
v 1-800-IBM-4YOU (426-4968) for DB2 marketing and sales
Product information
Information regarding DB2 Universal Database products is available by telephone
or by the World Wide Web at [Link]
This site contains the latest information on the technical library, ordering books,
product downloads, newsgroups, FixPaks, news, and links to web resources.
If you live in the U.S.A., then you can call one of the following numbers:
v 1-800-IBM-CALL (1-800-426-2255) to order products or to obtain general
information.
v 1-800-879-2755 to order publications.
For information on how to contact IBM outside of the United States, go to the IBM
Worldwide page at [Link]/planetwide
Printed in USA
SC09-4821-01
Spine information:
® ™
IBM DB2 Universal Database Administration Guide: Performance Version 8.2