0% found this document useful (0 votes)
6 views202 pages

IBMSIQData Server Admin Guide 76021

The IBM StoredIQ Data Server Administration Guide provides comprehensive instructions for managing administrative tasks related to IBM StoredIQ, including system configuration, volume management, and data harvesting. It outlines the user interface, system administration, and various settings for configuring the data server. Additionally, it includes information on accessing audits, logs, and deploying customized web services.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views202 pages

IBMSIQData Server Admin Guide 76021

The IBM StoredIQ Data Server Administration Guide provides comprehensive instructions for managing administrative tasks related to IBM StoredIQ, including system configuration, volume management, and data harvesting. It outlines the user interface, system administration, and various settings for configuring the data server. Additionally, it includes information on accessing audits, logs, and deploying customized web services.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

IBM StoredIQ

Data Server Administration Guide

IBM
Note
Before using this information and the product it supports, read the information in Notices.

This edition applies to Version [Link] of product number 5724M86 and to all subsequent releases and modifications
until otherwise indicated in new editions.
© Copyright International Business Machines Corporation 2001, 2020.
US Government Users Restricted Rights – Use, duplication or disclosure restricted by GSA ADP Schedule Contract with
IBM Corp.
Contents

Tables................................................................................................................... v

About this publication......................................................................................... viii


IBM StoredIQ product library....................................................................................................................viii
Contacting IBM StoredIQ customer support............................................................................................ viii

Overview of IBM StoredIQ Data Server...................................................................1

IBM StoredIQ Data Server user interface................................................................2


Logging in and out of IBM StoredIQ Data Server........................................................................................ 2
Navigation within IBM StoredIQ Data Server..............................................................................................2
Web interface icons..................................................................................................................................... 4

System administration...........................................................................................6
Checking the system's status...................................................................................................................... 6
Restarting the system.................................................................................................................................. 7

Configuration of IBM StoredIQ system settings...................................................... 9


Configuring DA Gateway settings................................................................................................................ 9
Configuring network settings.....................................................................................................................10
Configuring mail settings........................................................................................................................... 11
Configuring SNMP settings........................................................................................................................ 11
Configuring notification from IBM StoredIQ............................................................................................. 12
Configuration of multi-language settings..................................................................................................12
Setting the system time and date............................................................................................................. 14
Backing up the data server configuration................................................................................................. 14
Restoring system backups........................................................................................................................ 14
Backing up the IBM StoredIQ image.........................................................................................................15
Managing LDAP connections on the data server.......................................................................................16
Managing IBM StoredIQ Data Server administrative accounts................................................................ 17
Importing IBM Notes ID files ....................................................................................................................18
Configuring harvester settings.................................................................................................................. 19
Configuring full-text index settings........................................................................................................... 20
Specifying data object types......................................................................................................................22
Configuring audit settings..........................................................................................................................23
Configuring hash settings.......................................................................................................................... 23

Volumes and data sources................................................................................... 25


Volume indexing........................................................................................................................................ 25
Server platform configuration................................................................................................................... 25
Creating primary volumes..........................................................................................................................30
Creating retention volumes....................................................................................................................... 70
Creating system volumes.......................................................................................................................... 72
Export and import of volume data.............................................................................................................73
Creating discovery export volumes........................................................................................................... 75
Deleting volumes....................................................................................................................................... 76
Action limitations for volume types...........................................................................................................77
Volume limitations for migrations............................................................................................................. 80

iii
Data harvesting................................................................................................... 81
Harvest of properties and libraries............................................................................................................82
Lightweight harvest parameter settings................................................................................................... 82
Pausing and resuming harvests.................................................................................................................84
Monitoring harvests................................................................................................................................... 84

Job configuration................................................................................................ 89
Creating a job............................................................................................................................................. 89
Starting a job.............................................................................................................................................. 90
Monitoring processing............................................................................................................................... 91

Desktop collection............................................................................................... 94
IBM StoredIQ Desktop Data Collector client prerequisites..................................................................... 94
Downloading the IBM StoredIQ Desktop Data Collector installer .......................................................... 95
IBM StoredIQ Desktop Data Collector installation methods................................................................... 95
Configuring desktop collection settings....................................................................................................96
Configuring desktop collection..................................................................................................................97
Special considerations for desktop delete actions...................................................................................98

Audits and logs....................................................................................................99


Harvest audits............................................................................................................................................ 99
Import audits........................................................................................................................................... 102
Event logs.................................................................................................................................................103
Policy audits.............................................................................................................................................127

Deploying SharePoint customized web services..................................................133

Configuring administration knobs...................................................................... 134

Supported file types.......................................................................................... 135


Supported file types by name................................................................................................................. 135
Supported file types by category............................................................................................................ 143
SharePoint supported file types..............................................................................................................151

Supported server platforms and protocols..........................................................154

Event log messages........................................................................................... 162

Notices..............................................................................................................185
Trademarks..............................................................................................................................................186
Terms and conditions for product documentation................................................................................. 187
IBM Online Privacy Statement................................................................................................................ 187

Index................................................................................................................ 189

iv
Tables

1. IBM StoredIQ primary tabs........................................................................................................................... 2

2. Dashboard subtab settings and descriptions............................................................................................... 3

3. Configuration settings and descriptions....................................................................................................... 3

4. IBM StoredIQ dashboard icons.....................................................................................................................4

5. IBM StoredIQ Folders icons.......................................................................................................................... 4

6. System and Application configuration options.............................................................................................9

7. Supported languages.................................................................................................................................. 12

8. CIFS/SMB or SMB2 (Windows platforms) primary volumes......................................................................31

9. NFS v2 and v3 primary volumes................................................................................................................. 33

10. Exchange primary volumes.......................................................................................................................35

11. SharePoint primary volumes.....................................................................................................................38

12. Documentum primary volumes................................................................................................................ 42

13. Domino primary volumes..........................................................................................................................43

14. FileNet primary volumes...........................................................................................................................46

15. NewsGator primary volumes.................................................................................................................... 47

16. Livelink primary volumes.......................................................................................................................... 48

17. Jive primary volumes................................................................................................................................ 50

18. Chatter primary volumes.......................................................................................................................... 52

19. IBM Content Manager primary volumes...................................................................................................54

20. CMIS primary volumes..............................................................................................................................56

21. HDFS primary volumes............................................................................................................................. 57

22. Connections primary volumes.................................................................................................................. 59

23. SharePoint volumes as primary volumes: Fields and examples............................................................. 65

v
24. CIFS (Windows platforms) retention volumes......................................................................................... 70

25. NFS v3 retention volumes.........................................................................................................................71

26. System volume fields, descriptions, and applicable volume types.........................................................72

27. Discovery export volumes: fields, required actions, and applicable volume types................................ 76

28. Action limitations for volume types..........................................................................................................77

29. Predefined jobs......................................................................................................................................... 89

30. Harvest/Volume cache details: Fields, descriptions, and values............................................................ 91

31. Discovery export cache details: Fields, descriptions, and values........................................................... 91

32. Harvest audit by volume: Fields and descriptions................................................................................... 99

33. Harvest audit by time: Fields and descriptions........................................................................................99

34. Harvest overview summary options: Fields and descriptions...............................................................100

35. Harvest overview results options: Fields and descriptions................................................................... 100

36. Harvest overview detailed results: Fields and descriptions..................................................................100

37. Imports by volumes details: Fields and descriptions............................................................................102

38. ERROR event log messages....................................................................................................................104

39. INFO event log messages....................................................................................................................... 112

40. WARN event log messages..................................................................................................................... 120

41. Policy audit by name: Fields and descriptions.......................................................................................127

42. Policy audit by volume: Fields and descriptions....................................................................................127

43. Policy audit by time: Fields and descriptions.........................................................................................127

44. Policy audit by discovery exports: Fields and descriptions................................................................... 128

45. Discovery export runs by discovery export: Fields and descriptions.................................................... 128

46. Policy audit details: Fields and descriptions..........................................................................................129

47. Policy audit execution details: Fields and descriptions.........................................................................130

48. Policy audit data object details: Fields and descriptions...................................................................... 130

vi
49. Types of and reasons for policy audit messages................................................................................... 131

50. Admin knobs........................................................................................................................................... 134

51. Supported file types by name.................................................................................................................135

52. Supported archive file types by category...............................................................................................143

53. Supported CAD file types by category....................................................................................................144

54. Supported database file types by category........................................................................................... 144

55. Supported email file types by category..................................................................................................144

56. Supported graphic file types by category...............................................................................................145

57. Supported multimedia file types by category........................................................................................ 147

58. Supported presentation file types by category......................................................................................147

59. Supported spreadsheet file types by category...................................................................................... 148

60. Supported system file types by category...............................................................................................149

61. Supported text and markup file types by category................................................................................149

62. Supported word-processing file types by category............................................................................... 149

63. Supported SharePoint object types........................................................................................................151

64. Attribute summary..................................................................................................................................152

65. ERROR event log messages....................................................................................................................162

66. INFO event log messages....................................................................................................................... 170

67. WARN event log messages..................................................................................................................... 178

vii
About this publication
IBM® StoredIQ® Data Server Administration Guide describes how to manage the administrative tasks such
as administering appliances, configuring IBM StoredIQ, or creating volumes and data sources.

IBM StoredIQ product library


The following documents are available in the IBM StoredIQ product library.
• IBM StoredIQ Overview Guide
• IBM StoredIQ Deployment and Configuration Guide
• IBM StoredIQ Data Server Administration Guide
• IBM StoredIQ Administrator Administration Guide
• IBM StoredIQ Data Workbench User Guide
• IBM StoredIQ Cognitive Data Assessment User Guide
• IBM StoredIQ Insights User Guide
• IBM StoredIQ Integration Guide
The most current version of the product documentation can always be found online: https://
[Link]/support/knowledgecenter/en/SSSHEC_7.6.0/welcome/[Link]

Contacting IBM StoredIQ customer support


For IBM StoredIQ technical support or to learn about available service options, contact IBM StoredIQ
customer support at this phone number:
• 1-866-227-2068
Or, see the Contact IBM web site at [Link]

IBM Knowledge Center


The IBM StoredIQ documentation is available in IBM Knowledge Center.

Contacting IBM
For general inquiries, call 800-IBM-4YOU (800-426-4968). To contact IBM customer service in the
United States or Canada, call 1-800-IBM-SERV (1-800-426-7378).
For more information about how to contact IBM, including TTY service, see the Contact IBM website at
[Link]

viii IBM StoredIQ: Data Server Administration Guide


Overview of IBM StoredIQ Data Server
IBM StoredIQ Data Server provides access to data-server functions. It allows administrators to configure
system and application settings, manage volumes, administer harvests, configure jobs and desktop
collection, manage folders, access audits and error logs, and deploy customized settings.
The following list provides an overview of the tasks that the administrator can accomplish with IBM
StoredIQ Data Server:
Configuring system and application settings
This includes these tasks:
• Configure the DA gateway.
• View and modify network settings, including host name, IP address, NIS domain membership, and
use.
• View and modify settings to enable the generation of email notification messages.
• Configure SNMP servers and communities.
• Manage notifications for system and application events.
• View and modify date and time settings for IBM StoredIQ.
• Set backup configurations.
• Manage LDAP connections.
• Manage users.
• Upload Lotus Notes user IDs so that encrypted NSF files can be imported into IBM StoredIQ.
• Specify directory patterns to exclude during harvests.
• Specify options for full-text indexing.
• View, add, and edit known data object types.
• View and edit settings for policy audit expiration and removal.
• Specify options for computing hash settings when harvesting.
• Specify options to configure the desktop collection service.
Managing volumes and data sources
A volume represents a data source or destination that is available on the network to the IBM StoredIQ
appliance, and they are an integral to IBM StoredIQ indexing your data. Only administrators can
define, configure, and add or remove volumes.
Administering harvests
Harvesting (or indexing) is the process or task by which IBM StoredIQ examines and classifies data in
your network. Within IBM StoredIQ Data Server, you can specify harvest configurations.
Configuring jobs
Configure and run jobs from IBM StoredIQ Data Server.
Configuring desktop collection
Set up desktop collection. IBM StoredIQ Desktop Data Collector enables desktops as a volume type
or data source, allowing them to be used just as other types of added data sources. It can collect PSTs
and compressed files.
Accessing audits and logs
View and download audit and log categories as needed.
Deploying customized web services
Deploy SharePoint custom web services.

© Copyright IBM Corp. 2001, 2020 1


IBM StoredIQ Data Server user interface
Describes the IBM StoredIQ Data Server web interface and outlines the features within each tab.
References to sections where you can find additional information on each topic are also provided.

Logging in and out of IBM StoredIQ Data Server


The system comes with a default administrative account for use with the IBM StoredIQ Data Server user
interface. Use this account for the initial setup of the data server.
1. Open the IBM StoredIQ Data Server user interface from a browser.
Ask your system administrator for the URL.
2. In the login window, enter your user name or email address and password.
To log in for the first time after the deployment, use the default administrative account: enter admin in
the email address field and admin in the password field. For security reasons, change the password
for the admin account as soon as possible, preferably the first time you log in. Also, create extra
administrators for routine administration.
3. Click Log In to enter the system.
Database Compactor hint: If someone tries to log in while the appliance is doing database
maintenance, the administrator can override the maintenance procedure and use the system. For
more information, see “Job configuration” on page 89.
4. To log out of the application, click User Account in the upper right corner, and then click the Log out
link.

Navigation within IBM StoredIQ Data Server


The primary tabs and subtabs found within the user interface provide you with the access to data server
functionality.

Primary IBM StoredIQ Data Server tabs


IBM StoredIQ users do most tasks with the web interface. The menu bar at the top of the interface
contains three primary tabs that are described in this table.

Table 1. IBM StoredIQ primary tabs


Tab name Description
Administration Allows Administrators to do various configurations on these subtabs:
Dashboard, Data Sources, and Configuration.
Folders Create folders and jobs; run jobs.
Audit Examines a comprehensive history of all harvests, run policies, imports, and
event logs.

Administration tab
The Administration tab includes these subtabs: Dashboard, Data Sources, and Configuration.
• Dashboard: The Dashboard subtab provides an overview of the system's current, ongoing, and
previous processes and its status. This table describes administrator-level features and descriptions.
• Data sources: The Data sources subtab is where administrators define servers and volumes. They can
be places that are indexed or copied to. Various server types and volumes can be configured for use in

2 IBM StoredIQ: Data Server Administration Guide


managing data. Administrators can add FileNet servers through the Specify servers area. Volumes are
configured and imported in the Specify volumes section.
• Configuration: The administrator configures system and application settings for IBM StoredIQ through
the Configuration subtab.

Table 2. Dashboard subtab settings and descriptions


Dashboard
setting Description
Page refresh Choose from 30-second, 60-second, or 90-second intervals to refresh the page.
Today's job View a list of jobs that are scheduled for that day with links to the job's summary.
schedule
System View a summary of system details, including system data objects, contained data
summary objects, volumes, and the dates of the last completed harvest.
Jobs in View details of each job step as it is running, including estimated time to completion,
progress average speed, total system and contained objects that are encountered, harvest
exceptions, and binary processing information.
Harvest Review the performance over the last hour for all harvests.
statistics
Event log Review the last 500 events or download the entire event log for the current date or
previous dates.
Appliance Provides a status view of the appliance. Restart the appliance through the about
status appliance link. View cache details for volumes and discovery exports.

Table 3. Configuration settings and descriptions


Configuration
setting Description
System • DA Gateway settings: Configure the DA Gateway host or IP address.
• Network settings: Configure the private and public network interfaces.
• Mail server settings: Configure what mail server to use and how often to send email.
• SNMP settings: Configure Simple Network Management Protocol (SNMP) servers
and communities.
• System time and date: Set the system time and date on the appliance.
• Backup configuration: Back up the system configuration of the server to the IBM
StoredIQ gateway server.
• Manage LDAP connections: Add, edit, and remove LDAP connections.
• Manage users: Add, remove, and edit users.
• Lotus Notes user administration: Add a Lotus Notes User.

IBM StoredIQ Data Server user interface 3


Table 3. Configuration settings and descriptions (continued)
Configuration
setting Description
Application • Harvester settings: Set basic parameters and limits, data object extensions and
directories to skip, and reasons to run binary processing.
• Full-text settings: Set full-text search limits for length of word and numbers and edit
stop words.
• Data object types: Set the object types that appear in the disk use by data object
type report.
• Audit settings: Configure how long and how many audits are kept.
• Hash settings: Configure whether to compute a hash when harvesting and which
hash.
• Desktop settings: Configure the desktop collection service.

Folders tab
Within the Folders tab, any type of user can create and manage application objects.

Audit tab
IBM StoredIQ audit feature allows Administrators to review all actions that are taken with the data server,
including reviewing harvests and examining the results of actions.

Web interface icons


The following tables describe the icons that are used throughout IBM StoredIQ web interface.

IBM StoredIQ icons

Table 4. IBM StoredIQ dashboard icons


Dashboard icon Description
User account The User account icon accesses your user account, provides information about
version and system times, and logs you out of the system. For more information, see
Logging In and Out of the System.
Inbox The inbox link provides you the access to the PDF audit reports.
Help Clicking the Help icon loads IBM StoredIQ technical documentation in a separate
browser window. By default, the technical documentation is loaded as HTML help.

Folders icons

Table 5. IBM StoredIQ Folders icons


Folder icon Description
New Use the New icon to add jobs and folders.
Action Use the Action icon to act on workspace objects, including the ability to move, and
delete jobs and folders.
Job Jobs tasks such as harvesting are either a step or a series of steps. For more
information, see “Job configuration” on page 89.

4 IBM StoredIQ: Data Server Administration Guide


Table 5. IBM StoredIQ Folders icons (continued)
Folder icon Description
Folder Folders are a container object that can be accessed and used by administrators. For
more information, see “Job configuration” on page 89.
Folder Up Folders are a container object that can be accessed and used by administrators. By
default, you view the contents of the Workspace folder; however, by clicking this icon,
you move to the parent folder in the structure. For more information, see “Job
configuration” on page 89.

Audit icons
No specialized icons are used on the Audit tab.

IBM StoredIQ Data Server user interface 5


System administration
System administration entails checking the system's status, restarting the appliance, backing up the IBM
StoredIQ image, and backing up or restoring the data server configuration.

Checking the system's status


You can check the system's status for information about various appliance details.
1. In IBM StoredIQ Data Server, go to Administration > Dashboard > Appliance status.
2. Click About appliance to open the Appliance details page. The Appliance details page shows the
following information:
• Node
• Harvester processes
• Status
• Software version
• View details link
3. Click the View details link for the controller. The following table defines appliance details and
describes the data that is provided for the node.
Option Description
View Shows software version and details of harvester processes running on the
appliance controller for the appliance.
details
Application Shows a list of all services and status, which includes this information:
services
• Service: the name of each service on the appliance component
• PID: the process ID associated with each service
• Current memory (MB): the memory that is being used by each service
• Total memory (MB): total memory that is being used by each service and all child
services
• Processor percentage: the percentage of processor usage for each service. This
value is zero when a service is idle.
• Status: the status of each service. Status messages include Running, Stopped,
Error, Initializing, Unknown.

System Shows a list of basic system information details and memory usage statistics.
services System information includes this information:
• System time: current time on the appliance component
• GMT Offset: the amount of variance between the system time and GMT.
• Time up: the time the appliance component was running since the last restart, in
days, hours, minutes, and seconds
• System processes: the total number of processes that are running on the node
• Number of processors: the number of processors in use on the component
• Load Average (1 Minute): the average load for system processes during a 1-
minute interval
• Load Average (5 Minutes): the average load for system processes during a 5-
minute interval

6 IBM StoredIQ: Data Server Administration Guide


Option Description

• Load Average (10 Minutes): the average load for system processes during a 10-
minute interval
Memory details include this information:
• Total: total physical memory on the appliance component
• In use: how much physical memory is in use
• Free: how much physical memory is free
• Cached: amount of memory that is allocated to disk cache
• Buffered: the amount of physical memory that is used for file buffers
• Swap total: the total amount of swap space available (in use plus free)
• System services Swap in use: the total amount of swap space that is used
• Swap free: the total amount of swap space free
• Database Connections
• Configured
• Active
• Idle
• Network interfaces
• Up or down status for each interface

Storage Storage information for a controller includes this information:


• Volume
• Total space
• Used space
• Percentage

Controller and Indicator lights show component status


compute node
• Green light: Running
status
• Yellow light: The node is functional but is in the process of rebuilding;
performance can be degraded during this time. Note: The rebuild progresses
faster if the system is not being used.
• Red light: not running
Expand the node to obtain details of the appliance component by clicking the
image.

Restarting the system


Restart the system periodically.
The web application is temporarily unavailable if you restart it. Additionally, whenever volume definitions
are edited or modified, you must restart the system.
1. In IBM StoredIQ Data Server, select Administration > Dashboard > Appliance status. On the
Appliance status page, you have two options:
• Click the Controller link.
• Click About Appliance.
The Restart services and Reboot buttons are available on the Appliance details page and on each tab
of the Controller details page.

System administration 7
2. Click either of these options:
• Restart services: Restart all system services that are running on the node.
• Reboot: Restart the components.

8 IBM StoredIQ: Data Server Administration Guide


Configuration of IBM StoredIQ system settings
An administrator can modify the Application and Network areas to configure IBM StoredIQ.
The Configuration subtab (Administration > Configuration) is divided into System and Application
sections.

Table 6. System and Application configuration options


Section Configuration options
System • Configure the DA gateway.
• View and modify network settings, including host name, IP address, NIS
domain membership, and use.
• View and modify settings to enable the generation of email notification
messages.
• Configure SNMP servers and communities.
• Manage notifications for system and application events.
• View and modify date and time settings for IBM StoredIQ.
• Back up the data server configuration.
• Configure LDAP connections.
• Manage users.
• Upload Lotus Notes user IDs so that encrypted NSF files can be imported
into IBM StoredIQ.

Application • Specify directory patterns to exclude during harvests.


• Specify options for full-text indexing.
• View, add, and edit known data object types.
• View and edit settings for policy audit expiration and removal.
• Specify options for computing hash settings when harvesting.
• Specify options to configure the desktop collection service.

Configuring DA Gateway settings


View or modify the DA gateway settings. These settings are part of the general system-configuration
options for the data server.
1. Go to Administration > Configuration > DA Gateway settings.
2. If secure gateway communication (via stunnel) was enabled during deployment, the Host field
displays [Link]. If secured gateway communication was not enabled during deployment, the Host
field displays the IP address configured during deployment. You can update the IP address or enter
the host name fully qualified domain name of the StoredIQ gateway server instead.
For example, enter [Link] or [Link].
3. The Node name field shows the name of the data server that you assigned during installation.
Change as required.
4. If you changed any of the settings, restart services in either of the following ways:
• Go to the data server dashboard, click About Appliance and then click Restart Services.
• Using an SSH tool, log in to the data server VM as root and then run this command: service
deepfiler restart.

© Copyright IBM Corp. 2001, 2020 9


Configuring network settings
Describes how to configure the network settings that are required to operate IBM StoredIQ Data Server.
1. Go to Administration > Configuration > System > Network settings.
2. Click Controller Settings. Set or modify the following Primary Network Interface options.
• IP type: Set to static or dynamic. If it is set to dynamic, the IP address, Netmask, and Default
Gateway fields are disabled.
• IP address: Enter the IP address.
• Netmask: Enter the network mask of the IP address.
• Default gateway: Enter the IP address of the default gateway.
• Hostname: Enter the fully qualified domain name that is assigned to the appliance.
• Ethernet speed: Select the Ethernet speed.
• Separate network for file/email servers: Specify the additional subnet for accessing file/email
servers. If you are using the web application from one subnet and harvesting from another subnet,
select this checkbox.
For information about enabling or disabling the ports, see the topic about default open ports in the
deployment guide.
Restart the system for any primary network interface changes to take effect. See Restarting and
Rebooting the Appliance.
3. In Controller Settings, set or modify the following DNS Settings options.
• Nameserver 1: Set the IP address of the primary DNS server for name resolution.
• Nameserver 2: Set the IP address of the secondary DNS server for name resolution.
• Nameserver 3: Set the IP address of the tertiary DNS server for name resolution.
DNS settings take effect after they are saved. Changes to the server's IP address take effect
immediately. Because the server has a new IP address, you must reflect this new address in the
browser address line before next step.
4. Click OK.
5. Click Server name resolution. Set these options for the data server:
a) CIFS file server name resolution: These settings take effect upon saving.
• LMHOSTS: Enter the IP host name format.
• WINS Server: Enter the name of the WINS server.
b) NIS (for NFS): These settings take effect upon saving.
• Use NIS: Select this box to enable NIS to set UID/GID to friendly name resolution in an NFS
environment.
• NIS Domain: Specify the NIS domain.
• Broadcast for server on local network: If the NIS domain server is on the local network and can
be discovered by broadcasting, select this box. This option does not work if the NIS domain
server is on another subnet.
• Specify NIS server: If not using broadcast, specify the IP address of the NIS domain server here.
c) Doc broker settings (for Documentum)
• Enter Host for doc broker.
6. Click OK.

10 IBM StoredIQ: Data Server Administration Guide


Configuring mail settings
Mail settings can be configured as part of system configuration options.
1. Go to Administration > Configuration > System > Mail Server settings.
2. In Mail server, enter the fully qualified host name of the SMTP mail server.
3. In From address, enter a valid sender address. If the sender is invalid, some mail servers reject email.
A sender address also simplifies the process of filtering email notifications that are based on the
sender's email.
4. Click OK to save changes.

Configuring SNMP settings


You can configure the system to make Object Identifier (OID) values available to Simple Network
Management Protocol (SNMP) client applications. At the same time, you can receive status information or
messages about system events in a designated trap. For information about environmental circumstances
that are monitored by IBM StoredIQ, see the following table.
1. Go to Administration > Configuration > System > SNMP settings.
2. To make OID values available to SNMP client applications, in the Appliance Public MIB area:
a) Select the Enabled checkbox to make the MIB available, that is, to open port 161 on the controller.
b) In the Community field, enter the community string that the SNMP clients use to connect to the
SNMP server.
c) To view the MIB, click Download Appliance MIB. This document provides the MIB definition,
which can be provided to an SNMP client application.
3. To capture messages that contain status information in the Trap destination area:
a) In the Host field, enter the common name or IP address for the host.
b) In the Port field, enter the port number. Port number 162 is the default.
c) In the Community field, enter the SNMP community name.
4. To modify the frequency of notifications, complete these fields in the Environmental trap delivery
area:
a) Send environmental traps only every __ minutes.
b) Send environmental traps again after __ minutes.
5. Click OK. Environmental traps that are monitored by IBM StoredIQ are described in this table.
Option Description
siqConsoleLogLineTrap A straight conversion of a console log line into a trap. It uses these
parameters: messageSource, messageID, severity, messageText.
siqRaidControllerTrap Sent when the RAID controller status is anything but normal. Refer to the
MIB for status code information. It uses this parameter: nodeNum.
siqRaidDiskTrap Sent when any attached raid disk's status is anything but OK. It uses this
parameter: nodeNum.
siqBbuTrap Battery Backup Unit (BBU) error on the RAID controller detected. It uses
this parameter: nodeNum.
siqCacheBitTrap Caching indicator for RAID array is off. It uses this parameter: nodeNum.
siqNetworkTrap Network interface is not UP when it must be. It uses this parameter:
nodeNum.
siqDbConnTrap Delivered when the active Postgres connection percentage exceeds an
acceptable threshold. It uses this parameter: nodeNum.

Configuration of IBM StoredIQ system settings 11


Option Description
siqFreeMemTrap Delivered when available memory falls too low. It uses this parameter:
nodeNum.
siqSwapUseTrap Sent when swap use exceeds an acceptable threshold. Often indicates
memory leakage. It uses this parameter: nodeNum.
siqCpuTrap Sent when processor load averages are too high. It uses this parameter:
nodeNum.
siqTzMismatchTrap Sent when the time zone offset of a node does not match the time zone
offset of the controller. It uses this parameter: nodeNum.

Configuring notification from IBM StoredIQ


You can configure the system to notify you using email or SNMP when certain events occur.
For a list of events that can be configured, see “Event log messages” on page 104.
1. Go to Administration > Configuration > System > Manage notifications.
2. Click Create a notification.
3. In the Event number: field, search for events by clicking Browse or by typing the event number or
sample message into the field.
4. Select the event level by clicking the ERROR, WARN, or INFO link.
5. Scroll through the list, and select each event by clicking it. The selected events appear in the create
notification window. To delete an event, click the delete icon to the right of the event.
6. In the Destination: field, select the method of notification: SNMP, or Email address, or both. If you
choose email address, enter one or more addresses in the Email address field. If you choose SNMP,
the messages are sent to the trap host identified in the SNMP settings window, with a trap type of
siqConsoleLogLineTrap.
7. Click OK.
8. To delete an item from the Manage notifications window, select the checkbox next to the event, and
then click Delete.
You can also request a notification for a specific event from the dashboard’s event log. Click the
Subscribe link next to any error message and a prepopulated edit notification screen that contains the
event is provided.

Configuration of multi-language settings


The following table lists the languages that are supported by the IBM StoredIQ.
This list does not apply to IBM StoredIQ Cognitive Data Assessment. With CDA, the only supported
language is English.

Table 7. Supported languages


Language Code Lemmas Stop words
Arabic ar X
Catalan ca
Chinese zh X
Czech cs X
Danish da X

12 IBM StoredIQ: Data Server Administration Guide


Table 7. Supported languages (continued)
Language Code Lemmas Stop words
Dutch nl X
English en X X
Finnish fi X
French fr X X
German de X X
Greek el X
Hebrew he X
Hungarian hu
Icelandic is
Italian it X
Japanese ja X
Korean ko X
Malay ms
Norwegian (Bokmal) nb X
Norwegian (Nynorsk) nn X
Polish pl X
Portuguese pt X X
Romanian ro
Russian ru X
Spanish es X X
Swedish sv X
Thai th X
Turkish tr X
Vietnamese vi

By default, English is the only language that Multi-language Support identifies during a harvest and it is
also the default search language. Both the identified language (or languages) and the search default
language can be changed in the [Link] file on the data server. You can find this
properties file in the /usr/local/tomcat/webapps/storediq/WEB-INF/classes directory on each
data server. All versions of the [Link] file must be kept in sync across all data
servers for searches to be consistent and correct.
To change the language that the harvester can identify, use the [Link] field, which is
the second-to-last line of the file: [Link] = en,fr,de,pt. The first language in
the list is the default language, which is assigned to a document whose language cannot be identified.
To change the default search language, use the [Link] field, which is the last line of the
file. The search language is used to determine which language's rules, that is, stop words, lemmas,
character normalization, apply in a search. Only one language can be set as the default for search.
However, the default language can be manually overwritten in a full-text search: lang:de[umkämpft
großteils]

Configuration of IBM StoredIQ system settings 13


After you change this property file, you must restart the data server and reharvest the volumes that are to
be searched. If the data server is the DataServer - Distributed type, run the following command on the
data server after restarting it:

/etc/deepfile/dataserver/[Link]

Setting the system time and date


The system's time and date can be modified as needed.
A system restart is required for any changes that are made to the system time and date. See Restarting
and Rebooting the Appliance.
1. Go to Administration > Configuration > System > System time and date.
2. Enter the current date and time.
3. Select the appropriate time zone for your location.
4. Enable Use NTP to set system time to use an NTP server to automatically set the system date and
time for the data server.
If NTP is used to set the system time, then the time and date fields set automatically. However, you
must specify the time zone.
5. Enter the name or IP address of the NTP server.
6. Click OK to save changes.

Backing up the data server configuration


To prepare for disaster recovery, you can back up the system configuration of an IBM StoredIQ data
server to an IBM StoredIQ gateway server. This process backs up volume definitions, discovery export
records, and data-server settings. It does not back up infosets, data maps, or indexes. The preferred
method is to take a snapshot of the virtual machine to use as backup.
The gateway must be configured manually to support this backup.
1. Using an SSH tool, log in to the gateway as root.
2. At the command prompt, enter su util.
3. In the Appliance Manager utility, select Appliance Tools and OK, then press Enter.
4. Select Enable NFS Share for Gateway and press Enter.
5. Select Appliance Tools > OK > Enable NFS Share for Gateway.
A dialog box appears, stating that the system is checking exports.
6. Log in to the data server as an administrator.
7. Start the system backup. Go to Administration > Configuration > System > Backup configuration.
Click Start backup.
8. Check the event log. Go to Navigate to the Administration > Dashboard page and check the Event log
for the status information.
You can also examine the files created by the backup procedure. To see these files, go to the /
deepfs/backup directory on the gateway server.

Restoring system configuration backups


To prepare for disaster recovery, you can back up the system configuration of a IBM StoredIQ data server
to a IBM StoredIQ gateway server. You can later restore a system's configuration.
1. Build a new data server, which serves as the restore point for the backup. When building this new data
server, use the same IP address and server name as the original data server. Verify that the original
server is shut down and not on the network.

14 IBM StoredIQ: Data Server Administration Guide


See the topics about configuring the data server and the data server gateway settings in the
deployment guide for the installation and configuration procedures.
2. Ensure the gateway still has its configured NFS mount.
a) Using an SSH tool, log in to the gateway as root.
b) At the command prompt, enter su util.
c) In the Appliance Manager utility, select Appliance Tools and OK, then press Enter.
d) Select Enable NFS Share for Gateway and press Enter.
A dialog box appears, stating that the system is checking exports.
3. Using an SSH tool, log in to the new data server as root.
4. At the command prompt, enter su util.
5. In the Appliance Manager utility, select Appliance Tools and OK, then press Enter.
6. Select Restore Configuration From Backups and press Enter.
7. Enter the gateway's server name or IP address, and then press Enter.
8. Provide the full system restore date, or leave the space empty in order to restore the most recent
system backup.
9. Enter Y and then press Enter to confirm the system's restoration.

Backing up the IBM StoredIQ image


Backing up the IBM StoredIQ images is a good method for disaster recovery. It is also a best practice
before you start any upgrades on your images. If you need to back up the IBM StoredIQ images, you must
complete the following steps.
An active IBM StoredIQ image must not be backed up by using VMWare VCenter or other product backup
utilities. If you do so, the data servers might hang and become unresponsive. Running a backup snapshot
on an active IBM StoredIQ image might result in transaction integrity issues.
To prepare for disaster recovery, another method is to back up the system configuration of the IBM
StoredIQ data server to an IBM StoredIQ gateway server. This type of backup is supported only for data
servers.
If a backup snapshot of IBM StoredIQ image is needed, follow these steps:
1. Stop services on all data servers and the gateway:
a) Log in to each data server and to the gateway as root.
b) To stop all IBM StoredIQ services, enter the following command:

service deepfiler stop

c) To stop the postgresql database service, enter the following command:

service postgresql stop

d) Log out.
Important: Wait 10 minutes after a harvest before you use this command to stop services.
2. Stop the IBM StoredIQ services on the application stack:
a) Log in to the application stack as siqadmin user.
Alternatively, you can log in as root user.
b) Enter the following command:

systemctl stop [Link]

c) Log out.

Configuration of IBM StoredIQ system settings 15


3. Contact the VMWare VCenter administrator to have a snapshot of the IBM StoredIQ image taken.
Confirm the work completion before you proceed to the next step.
4. Restart services on all data servers and the gateway:
a) Log in to each data server and to the gateway as root.
b) To restart all IBM StoredIQ services, enter the following command:

service deepfiler restart

c) To restart the postgresql database service, enter the following command:

service postgresql restart

5. Start the IBM StoredIQ services on the application stack:


a) Log in to the application stack as siqadmin user.
Alternatively, you can log in as root user.
b) Enter the following command:

systemctl start [Link]

Managing LDAP connections on the data server


Configure and manage connections to the LDAP server so that you can create administrative users who
use LDAP authentication when logging in to IBM StoredIQ Data Server.
You must be logged in to IBM StoredIQ Data Server.
To be able to create LDAP users in IBM StoredIQ Data Server, at least one LDAP connection must be
configured and defined as default connection for authentication.
1. On the Administration > Configuration page, click Manage LDAP connections.
2. Select one of the following options:
• To add connections:
a. Click Add connection. Then, provide the following connections details:

Parameter Value
LDAP server The URL of the LDAP server in the form of an IP address or the
FQDN (fully qualified domain name).
Principal The security principal for this connection in this format:
cn=common_name,ou=organizational_unit,dc=domain_component

If you configure this connection as default connection, the


principal must have admin privileges on the LDAP server and
must belong to one of the users that you want to add.

Password The principal's password.


b. Make this connection the default connection for LDAP authentication. One of the configured
connections must be set as default connection before you can add LDAP users. The default
connection is used when validating other LDAP users with the LDAP server. Therefore, the
security principal that you specify for this connection must have admin privileges on the LDAP
server and must belong to one of the users that you want to add.
c. Click OK.
d. Add further connections by repeating steps “2.a” on page 16 to “2.c” on page 16.

16 IBM StoredIQ: Data Server Administration Guide


• To edit a connection, click the respective entry and update the settings as required. Remember that
you must enter the password again to apply the changes.
• To delete a connection, click the respective entry and then click Delete on the connection details
window. You can delete only connections that are not in use.

Managing IBM StoredIQ Data Server administrative accounts


Account administration in IBM StoredIQ Data Server includes creating, modifying, and deleting
administrative accounts.
If you want to create users who authenticate through LDAP, configure at least a default LDAP connection
before you start adding LDAP users. For more information, see “Managing LDAP connections on the data
server” on page 16
IBM StoredIQ Data Server comes with a default administrative account that you can use for the initial
setup. The default system administrator is admin. The default password for this account is admin. For
security purposes, change the password as soon as possible. Also, create additional administrative
accounts on the data server for routine administration so that the actions can be audited. These accounts
allow access only to the IBM StoredIQ Data Server interface on data server where they were created.
Note: If someone tries to log in while the Database Compactor appliance is doing database maintenance,
the administrator can override the maintenance procedure and use the system. For more information, see
“Job configuration” on page 89.
1. In a browser window, enter the IP address or host name of IBM StoredIQ Data Server.
2. Log in with your IBM StoredIQ Data Server credentials.
To log in for the first time after the deployment, use the default administrative account. For regular
accounts, use the email address that is defined in your account settings to log in. Local users must
provide the password they configured in IBM StoredIQ Data Server. LDAP users are authenticated by
using the LDAP server and must therefore provide their LDAP password.
3. On the Administration > Configuration page, click Manage users.
You have the following options:
• Change the password of the default administrative account.
From the list, click The Administrator account, and then select Change the “admin” password.
This default administrative account is always a local account.
• Create an account.
Click Create new user. Provide the name and an email address, and select the authentication type
and the appropriate notification setting.
For a user with the authentication type LDAP, you must also provide the LDAP principal of the user
you want to add in this format:
cn=common_name,ou=organizational_unit,dc=domain_component
For details, see your LDAP documentation.
When you click OK, LDAP user information is verified with the LDAP server to make sure that the
user exists and that the attributes are valid. Therefore, at least a default LDAP connection must be
configured before you can create LDAP users.
Users with the authentication type Local must create and maintain a password for authenticating to
IBM StoredIQ Data Server. The welcome message that they receive provides instructions for
creating this password. LDAP users authenticate with their LDAP passwords.
• Edit an account.
In the list, click the user name of the account you want to edit. Then, click Edit user and change
settings as required. For an LDAP user, the information is verified with the LDAP server when you
click OK.

Configuration of IBM StoredIQ system settings 17


• Lock or unlock a local or LDAP user account.
In the list, click the user name of the account you want to lock or unlock. Then, click either Lock
account or Unlock account. A locked account is marked accordingly. An account also becomes
locked after three failed login attempts. The user cannot log in while the account is locked.
• Change your password.
This option is available only for local user accounts. Passwords of LDAP users must be changed by
using LDAP administration tools.
In the list, click your user name to open your account. Then, click Change password. Alternatively,
open your account by selecting Your account from the user menu in the navigation bar.
When you need to change your password while you are not logged in, you can click the Forgot your
password? link in the login window. You will then receive an email with instructions for creating a
new password. Again, this applies to local user accounts only.
• Reset other users' passwords.
This option is available only for local user accounts. Passwords of LDAP users must be reset by
using LDAP administration tools.
In the list, click the user name of the account. Then, click Reset password. The user receives an
email with the information that the password was reset and instructions for creating a new
password.
• Delete an account.
In the list, click the user name of the account. Then, click Delete.
When you delete an LDAP user, only the IBM StoredIQ Data Server account is deleted. The user
account on the LDAP server remains unchanged.

Importing IBM Notes ID files


IBM StoredIQ can decrypt, import, and process encrypted NSF files from IBM Lotus Domino. The feature
works by comparing a list of [Link] and key pairs that were imported into the system with the key values
that lock each encrypted container or email. When the correct match is found, the file is unlocked with the
key. After the emails or containers are unlocked, IBM StoredIQ analyzes and processes them in the usual
fashion.
These use cases are supported:
• Multiple unencrypted emails within a journaling database that was encrypted with a single [Link] key
• Multiple unencrypted emails in an encrypted NSF file
• Multiple encrypted emails within an unencrypted NSF file
• Multiple encrypted emails with the same or different [Link] keys, contained in an encrypted NSF file
• Encrypted emails from within a journaling database
1. On the primary data server, go to Administration > Configuration > Lotus Notes user administration.
2. Click Upload a Lotus user ID file.
a) In the dialog box that appears, click Browse, and go to a user ID file.
b) Enter the password that unlocks the selected file.
c) Enter a description for the file.
d) Click OK. Repeat until the keys for all encrypted items are uploaded. When the list is compiled, you
can add new entries to it.
e) To delete an item from the list, from the Registered Lotus users screen, select the checkbox next
to a user, and then click Delete. In the confirmation dialog that appears, click OK.
Note: After you upload user IDs, restart services.

18 IBM StoredIQ: Data Server Administration Guide


Configuring harvester settings
You can use several different harvester settings to fine-tune your index process.
1. Go to Administration > Configuration > Application > Harvester settings.
2. To configure Basic settings, follow these steps:
a) Harvester Processes: Select either Content processing or System metadata only.
b) Harvest miscellaneous email items: Select to harvest contacts, calendar appointments, notes,
and tasks from the Exchange server.
c) Harvest non-standard Exchange message classes: Select to harvest message classes that do not
represent standard Exchange email and miscellaneous items.
d) Include extended characters in object names: Select to allow extended characters to be included
in data object names during a harvest.
e) Determine whether data objects have NSRL digital signature: Select to check data objects for
NSRL digital signatures.
f) Enable parallel grazing: Select to harvest volumes that were already harvested and are going to be
reharvested.
If the harvest completes normally, parallelized grazing enables harvests to begin where they left off
when interrupted and to start at the beginning.
g) Index generated text: Select for the generated text that is extracted by OutsideIn to be indexed
and available for full-text search.
For sensitive data detection, especially in spreadsheet files, this option must be enabled.
3. Specify Skip Content processing.
In Data object extensions to be skipped, specify those file types that you want the harvest to ignore
by adding data object extensions to be skipped.
4. To configure Locations to ignore, enter each directory that must be skipped. IBM StoredIQ accepts
only one entry per line and that regular expressions can be used.
5. To configure Limits, follow these steps:
a) Maximum data object size: Specify the maximum data object size to be processed during a
harvest.
During a harvest, files that exceed the maximum data object size are not read. As a result, if full-
text/content processing is enabled for the volume, they are audited as skipped: Configured max.
object size. These objects still appear in the volume cluster along with all file system metadata.
Since they were not read, the hash is a hash of the file-path and size of the object, regardless of
what the hash settings are for the volume (full/partial/off).
b) Max entity values per entity: For any entity type (date, city, address and the like), the system
records, per data object, the number of values set in this field.
The values do not need to be unique. For example, if the maximum value is 1,000, and the
harvester collects 1,000 instances of the same date (8/15/2009) in a Word document, the system
stops counting dates. This setting applies to all user-defined expressions (keyword, regular
expression, scoped, and proximity) and all standard attributes.
c) Max entity values per data object: Across all entity types, the total (cumulative) number of values
that is collected from a data object during a harvest. A 0 in this field means "unlimited".
This setting applies to all user-defined expressions (key-word, regular expression, scoped, and
proximity) and all standard attributes.
6. Configure Binary Processing.
a) Run binary processing when text processing fails: Select this option to run binary processing.
The system runs further processes against content that failed in the harvesting. You can select
options for when to start this extended processing and how to scan content. Binary processing
does not search image file types such as .GIF and .JPG for text extraction.

Configuration of IBM StoredIQ system settings 19


b) Failure reasons to begin binary processing: Select the checkboxes of the options that define
when to start extended processing.
Binary processing can enact in extracting text from a file failure in these situations:
• when the format of the file is unknown to the system parameters;
• when the data object type is not supported by the harvester scan;
• when the data object format does not contain actual text.
c) Data object extensions: Set binary processing to process all data files or only files of entered
extensions. To add extensions, enter one per line without a period.
d) Text encoding: Set options for what data to scan and extract at the start of binary processing.
This extended processing can accept extended characters and UTF-16 and UTF-32 encoded
characters as text. The system searches UTF-16 and UTF-32 by default.
e) Minimums: Set the minimum required number of located, consecutive characters to begin
processing for text extraction.
For example, if you enter 4, the system begins text processing when four consecutive characters of
a particular select text encoding are found. This setting helps find and extract helpful data from the
binary processing, reducing the number of false positives.
7. Click OK.
Changes to harvester settings do not take effect until the appliance is rebooted or the application
services are restarted.

Optical character recognition processing


Optical character recognition (OCR) processing enables text extraction from graphic image files that are
stored inside archives where the Include content tagging and full-text index option is selected.
After content typing inside the IBM StoredIQ processing pipeline, enabling OCR processing routes the
following file types through an optical character recognition engine OCR to extract recognizable text.
• Windows or OS/2 bitmap (BMP)
• Tag image bitmap file (TIFF)
• Bitmap (CompuServe) (GIF)
• Portable Network Graphics (PNG)
• Joint Picture Experts Group (JPG)
The text that is extracted from image files is processed through the IBM StoredIQ pipeline in the same
manner as text extracted from other supported file types. Policies with a specific feature to write out
extracted text to a separate file for supported file types do so for image files while OCR processing is
enabled.
The OCR processing rate of image files is approximately 7-10 KB/sec per IBM StoredIQ harvester
process.

Configuring full-text index settings


Use the full-text index settings feature to customize your full-text index.
Before you configure or search the full-text index, consider the following situations:
• Full-text filters that contain words might not return all instances of those words: You can limit full-text
indexing for words that are based on their length. For example, if you choose to full-text index words
limited to 50 characters, then no words greater than 50 characters are indexed.
• Full-text filters that contain numbers might not return all instances of those numbers: This situation can
occur when number searches are configured as follows:

20 IBM StoredIQ: Data Server Administration Guide


– The length of numbers to full-text index was defined. If you configure the full-text filter to index
numbers with 3 digits or more and try to index the numbers 9, 99, 999, and the word stock, only the
number 999 and the word stock are indexed. The numbers 9 and 99 are not indexed.
– Number indexing in data objects that are limited by file extensions. For example, if you choose to full-
text index the number 999 when it appears in data objects with the file extensions .XLS and .DOC,
then a full-text filter returns only those instances of the number 999 that exist in data objects with
the file extensions .XLS and .DOC. Although the number 999 can exist in other data objects that are
harvested, these data objects do not have the file extensions .XLS or .DOC.
1. Go to Administration > Configuration > Application > Full-text settings.
2. To configure Limits:
a) Do not limit the length of the words that are indexed: Select this option to have no limits on the
length of words that are indexed.
b) Limit the length of words indexed to___characters: Select this option to limit the length of words
that are indexed. Enter the maximum number of characters at which to index words. Words with
more characters than the specified amount are not indexed.
3. To configure Numbers:
• Do not include numbers in the full-text index: Select this option to have no indexed numbers.
This option is selected by default.
• Include numbers in the full-text index: Select this option to have numbers to be indexed.
• Include numbers in full-text index but limit them by: Select this option to have only certain
numbers indexed. Define these limits as follows:
– Number length: Include only numbers that are longer than ____ characters. Enter the number of
characters a number must contain to be indexed. The Number length feature indexes longer
numbers and ignores shorter numbers. By not indexing shorter numbers, such as one- and
two-character numbers, you can focus your filter on meaningful numbers. These numbers can
be account numbers, Social Security numbers, credit card numbers, license plate numbers, or
telephone numbers.
– Extensions: Index numbers that are based on the file extensions of the data objects in which
they appear. Select Limit numbers for all extensions to limit numbers in all file extensions to
the character limits set in Number length. Alternatively, select Limit numbers for these
extensions to limit the numbers that are selected in Numbers length only to data objects with
certain file extensions. Enter the file extensions one per line that must have limited number
indexing. Any data object with a file extension that is not listed has all indexed numbers.
For sensitive data detection, especially in spreadsheet files, make sure numbers are included in the
full-text index.
4. To configure Include word lemmas in index, select whether to identify and index the lexical forms of
words as well.
For example, employ is the lemma for words such as employed, employment, employs. If you use
lemmas and search for the word employed, IBM StoredIQ denotes any found instances of
employment, employ, employee, and so on, when it views the data object.
• Do not include word lemmas in index (faster indexing): By not indexing lemmas, data sources are
indexed slightly faster and the index size on disk is smaller.
• Include word lemmas in index (improved searching): By indexing lemmas, filter results can be
more accurate, although somewhat slower. Without lemmas, a filter for trade would need to be
written as trade, trades, trading, or traded to get the same effect, and even then a user might
miss an interesting variant.
5. Configure Stop words.
Stop words are common words that are found in data objects and are indexed like other words. This
allows users to find instances of these words where it matters most. A typical example would be a
search expression of 'to be or not to be' (the single quotation marks are a specific usage here).
Typically, IBM StoredIQ ignores stop words in search expressions, but because single quotation marks

Configuration of IBM StoredIQ system settings 21


as syntax elements, a user can find Shakespeare's "Hamlet." Indexing stop words slightly increases
the amount of required storage space, but relevant documents might be missed without these words
present in the index. By default, the following words are considered stop words for the English
language: a, an, and, are, as, at, be, but, by, for, if, in, into, is, it, no, not, of, on, or, such,
that, the, their, then, there, these, they, this, to, was, will, with.
To add a stop word, enter one word per line, without punctuation, which includes hyphens and
apostrophes.
Note: As of the IBM StoredIQ [Link] release, stop words on the configuration page are for the English
language only.
6. Select Enable OCR image processing to control at a global level whether Optical Character
Recognition (OCR) processing is attempted on image files, such as PNG, BMP, GIF, TIFF, and
images that are produced from scanned image PDF files. A scanned image PDF file is a PDF file with a
document that is scanned into it. Through OCR processing, the images are extracted from scanned
image PDF files and texts are extracted from the image files.
The quality of the text extraction relies on the resolution setting on the image files and images from
scanned image PDF files. Thus, the resolution setting must be at 300 dots per inch (DPI) or higher. For
text in images that is rotated, small font, or unclear text cannot be extracted.
If you select this option, you must restart services. See Restarting and Rebooting the Appliance.
7. Select Always process PDFs for images to control at a global level to extract text from scanned image
PDF files.
A scanned image PDF is a special type of PDF that is created by scanning a document into PDF and is
different from a normal PDF. A scanned image PDF contains one image per entire page and no other
elements such as plain text. In contrast, a normal PDF can contain a mix of plain-text elements,
embedded objects, and images per page. Text extraction from a scanned image PDF is processing
intensive and involves two steps:
a. Retrieving images from scanned image PDF
b. Extracting text from the retrieved images
To identify a PDF as a scanned image PDF and then extract text from it, you must select both the
Enable OCR image processing option and the Always process PDFs for images option. However, to
extract text from image files such as PNG, BMP, GIF, and TIFF, you need to select only the Enable OCR
image processing option (as described in step “6” on page 22) because only step b needs to be
performed on these files.
You can set a maximum number of images for processing by entering the respective count for the
Limit number of images in scanned image PDF to option. However, this setting does affect text
extraction only. The default value is zero, which means that text is extracted from all images that are
retrieved from scanned image PDFs.
If you select this option, you must restart services. See Restarting and Rebooting the Appliance.
8. Click OK.

Specifying data object types


On the Data object types page, you can add new data object types and view and edit known data object
types. These data objects appear in the Disk usage (by data object type) report. Currently, there are over
400 data object types available.
1. Go to Administration > Configuration > Application > Data object types.
2. In the add data object type section, enter one or more extensions to associate with the data object
type. These entries must be separates by spaces.
For example, enter doc txt xls.
3. Enter the name of the data object type to be used with the extension or extensions.
For example, enter Microsoft Word.

22 IBM StoredIQ: Data Server Administration Guide


4. Click Add to add the extension to the list.

Configuring audit settings


Audit settings can be configured to determine the number of days and number of policy audits to be kept
before they are deleted.
1. Go to Administration > Configuration > Application > Audit settings.
2. Specify the number of days to keep the policy audits before automatically deleting them.
3. Specify the maximum number of policy audits to keep before automatically deleting them.
4. Specify the file limit for drill-down in policy audits.
5. Click OK to save changes.

Configuring hash settings


Hashes are used to identify unique content. Configure the type of hash to compute when harvesting.
By default, IBM StoredIQ computes a SHA-1 hash for each object encountered during harvesting. If the
SHA-1 hash is based on the content of the files, it can be used to identify unique files (and duplicates).
For computing such a hash, document content must be fetched over the network even for harvests where
only file system metadata is collected. To avoid this, you can disable content based hashing for file
system metadata only indexing. This provides the fastest indexing rate at the expense of the ability to
identify unique content. In this case, the information used to compute the hash is based on volume and
object metadata.
If you change the hash settings between harvests, the next harvest uses the updated settings for any new
or modified documents. For example, you might not have content based hashes created initially, but
some time after the harvest completed you decide to enable content based hashing. In this case, a full-
text harvest (if the volume allows for that) generates regular content based hashes for all documents that
are indexed during the harvest.
The hash setting does not impact data object preview.
1. Go to Administration > Configuration > Application > Hash settings.
2. Determine whether you want to generate a content based hash.

• For content based hashes, leave the Compute data object hash option selected.
With this setting, content based hashes are generated as selected for full-text and metadata
harvests (see step “4” on page 24).
• For metadata based hashes, clear the Compute data object hash checkbox.
3. For creating a hash for email, select what email attributes are considered to compute the hash.
Email has characteristics that present a challenge when attempting to identify unique messages based
on a hash. Using a pure content based hash, it is likely that emails with identical user-visible content
do not share the same SHA-1 hash. Therefore, you can select from a set of attributes the ones to
contribute to the hash. By using specific fields to compute the email hash, an email located in a local
PST archive in a file system, for example, can be identified as a duplicate of a message in an Exchange
mailbox even though they are stored in completely different binary formats.
By default, the following information contributes to the hash:
• The information in the To, From, CC, and BCC attributes
• The email subject
• The content of the email body
• The content of any email attachments

Configuration of IBM StoredIQ system settings 23


The email hash selections operate independently from the data object hash settings; that is, a data
object can have a binary hash or an email hash, but not both.
4. For content based hashes, select whether you want to generate a full or a partial hash. This option is
not available if you cleared the Compute data object hash checkbox.
IBM StoredIQ offers two strategies for computing a content based hash. The default option is to read
the entire contents of each file as input to computing a SHA-1 hash for the file (full hash). If the
content of a file must be read to satisfy other content based index options (container processing or
full-text indexing), a full content based hash is always computed.
If you want only a file system metadata index with the ability to identify unique files, you have the
option to create a hash from parts of the file content (partial hash). With a partial hash , only a
maximum of 128 KB of a file's content is read to compute the hash. This minimizes the amount of data
read reducing the workload on the data source and network and effectively increasing the indexing
rate.
For a partial hash, up to four 32 KB blocks from each file are read to compute the hash. If a file is less
than 128 KB in size, the entire file content is evaluated. Content to compute the hash for files with a
size greater than 128 KB is read as follows:
• 1 x 32 KB block taken from the beginning of the file
• 2 x 32 KB blocks equally spaced between the beginning and end of the file
• 1 x 32 KB block taken from the end of the file
The resulting four 32 KB blocks are used as input to compute the hash. The partial hash might not be
appropriate for all use cases but might be sufficient for use cases such as storage management.
• For a full hash, leave Entire data object content (required for data object typing) selected.
IBM StoredIQ uses Oracle Outside In Technology filters to determine the object type based on
content and to extract additional metadata and text.
IBM StoredIQ implements its own support for text files, web archives (MHT), IBM Notes® email, and
EMC EmailXtender and SourceOne archives.
If a particular data object cannot be handled with the available text extraction methods, IBM
StoredIQ can selectively use binary processing to extract strings from a file. File processed in this
way have a binary processing attribute associated with them to allow the content to be filtered
based on this processing attribute. It can be useful to segregate these files because binary
processing can yield a high rate of false positives relative to other content extraction techniques.
You can configure binary processing in the harvester settings.
• For a partial hash, select Partial data object content.
5. Click OK.

24 IBM StoredIQ: Data Server Administration Guide


Volumes and data sources
Volumes or data sources are integral to IBM StoredIQ to index your data.
A volume represents a data source or destination that is available on the network to the IBM StoredIQ
appliance. A volume can be a disk partition or group of partitions that is available to network users as a
single designated drive or mount point. IBM StoredIQ volumes have the same function as partitions on a
hard disk drive. When you format the hard disk drive on your PC into drive partitions A, B, and C, you are
creating three partitions that function like three separate physical drives. Volumes behave the same way
that disk partitions on hard disk drive behave. You can set up three separate volumes that originate from
the same server or across many servers. Only administrators can define, configure, and add or remove
volumes to IBM StoredIQ.

Volume indexing
When you define volumes, you can determine the type and depth of index that is conducted.
Three levels of analysis are as follows.
• System metadata index. This level of analysis runs with each data collection cycle and provides only
system metadata for system data objects in its results. It is useful as a simple inventory of what data
objects are present in the volumes you defined and for monitoring resource constraints, such as file
size, or prohibited file types, such as the .MP3 files.
• System metadata plus containers. In a simple system metadata index, container data objects
(compressed files, PSTs, emails with attachments, and the like) are not included. This level of analysis
provides container-level metadata in addition to the system metadata for system data objects.
• Full-text and content tagging. This option provides the full local language analysis that yields the more
sophisticated entity tags. Naturally, completing a full-text index requires more system resources than a
metadata index. Users must carefully design their volume structure and harvests so that the maximum
benefit of sophisticated analytics is used, but not on resources that do not require them. Parameters
and limitations on “full-text” indexing are set when the system is configured.

Server platform configuration


IBM StoredIQ supports a wide variety of data sources, which can be added as volumes. Before you can
add volumes, you must configure the server platforms for the different volume types.
Each server type has prerequisite permissions and settings.

Defining server aliases


You can define server aliases for your volumes.
If your server naming conventions aren't very descriptive, you can assign an alias to the server for easier
reference. Another use case for an alias might be to allow for mapping several volumes to the same data
source such that volumes overlap one another. Exercise care when using server aliases for the second
use case because the same physical data object might appear on multiple volumes.
1. Navigate to Administration > Data sources > Specify servers and click Server aliases.
2. In the Add an alias section, enter a server and an appropriate alias, and click Add.
As server, specify the fully qualified domain name (FQDN) or IP address as appropriate for the server
type.
The information is displayed in the list of server aliases.
In IBM StoredIQ Data Server, you can now use the alias instead of the FQDN or IP address to map a
volume to the data source.

© Copyright IBM Corp. 2001, 2020 25


At any time, you can edit a server alias. However, in this case, any mappings in which the alias is used will
break. You cannot delete an alias that is in use.

Configuring Windows Share (CIFS)


Windows Share (CIFS) must be configured to harvest and run policies.
To harvest and run policies on volumes on Windows Share (CIFS) servers, the user must be in the backup
operator group on the Windows Share server that shows the shares on IBM StoredIQ and also needs to
have full control share-level permissions.
If the credential that is used for the connection to the server when the volume is created has the create
files permission, the modified date attribute on the shared folder on the CIFS server is updated. If
the volume is not used as a destination volume in a Copy or Move action and you don't want the
modified date attribute to be updated, remove the create files permission from the credential
used for the connection.

Configuring NFS
NFS must be configured to harvest and run policies.
• To harvest and run policies on NFS servers, you must enable root access on the NFS server that is
connected to IBM StoredIQ.

Configuration of Exchange servers


When you configure Exchange servers, you must consider various connections and permissions.
• Secure connection. If you want to connect to Exchange volumes over HTTPS, you can either select the
Use SSL checkbox or add port number 443 after the server name. If you choose the latter option, an
example is [Link]. In some cases, this secure connection can result in
some performance degradation due to SSL running large. If you enter the volume information without
the 443 suffix, the default connection is HTTP.
• Permissions for Exchange 2003. The following permissions must be set on the Exchange server to the
mailbox store or the mailboxes from which you harvest.
– Read
– Execute
– Read permissions
– List contents
– Read properties
– List object
– Receive as
• Permissions for Exchange 2007, 2010, 2013, and Online. The Full Access permissions must be granted
on the Exchange server for each mailbox from which you harvest.
The account that you specify in IBM StoredIQ when creating an Exchange Online volume must also have
the Read and manage mailbox permission for each mailbox to be harvested. You can grant this
permission in the Microsoft admin center.
• Deleted items. To harvest items that were deleted from the Exchange server, enable Exchange's
transport dumpster settings. For more information, see Microsoft® Exchange Server 2010
Administrator's Pocket Consultant. Configuration information is also available online at
[Link]. It applies only to on-premises versions of Exchange.
• Windows Authentication. For all on-premises versions, enable Integrated Windows Authentication on
each Exchange server.
• Public folders. To harvest public folders in Exchange, the Read Items privilege is required. It applies to
Exchange 2003 and 2007.

26 IBM StoredIQ: Data Server Administration Guide


• An Exchange 2013 service account must belong to an administrative group or groups granted the
following administrator roles:
– Mailbox Search
– ApplicationImpersonation
– Mail Recipients
– Mail Enabled Public Folders
– Public Folders
An Exchange Online service account must belong to an administrative group or groups granted the
following administrator roles, which are required as part of the Service account:
– Mailbox Search
– ApplicationImpersonation
– Mail Recipients
– Mail Enabled Public Folders
– MailboxSearchApplication
– Public Folders
Note: It is possible to create a new Exchange Admin Role specific to IBM StoredIQ that includes only
these roles.
The current Exchange Online authentication uses basic authentication over SSL. Volume credentials
that are supplied for Exchange Online are only as secure as the SSL session.
Note: Exchange Online connection uses claims-based authentication only. OAuth is not supported
currently.
Deleted items might persist because of Exchange Online's retention policies. Exchange Online is a
cloud-based service; items are deleted by an automated maintenance task. Items that are deleted
manually might persist until the automated job completes.

Enabling integrated Windows authentication on Exchange servers


Windows authentication can be integrated on Exchange servers.
1. From Microsoft Windows, log in to the Exchange Server.
2. Go to Administrative Tools > Internet Information Services (IIS) Manager.
3. Go to Internet Information Services > Name of Exchange Server > Web Sites > Default Web Site.
4. Right-click Default Web Site, and then click the Directory Security tab.
5. In the Authentication and access control pane, click Edit.
6. Select Properties. The Authentication Methods window appears.
7. In the Authentication access pane, select the Integrated Windows authentication check box.
8. Click OK.
9. Restart IIS services.

Improving performance for IIS 6.0 and Exchange 2003


Within IBM StoredIQ, performance can be improved for IIS 6.0 and Exchange 2003.
1. From Microsoft Windows, log on to the Exchange Server.
2. Go to Administrative Tools > Internet Information Services (IIS) Manager.
3. Select Internet Information Services > <Name of Exchange Server> > Web Sites > Application
Pools.
4. Right-click Application Pools and select Properties.
5. On the Performance tab, locate the Web Garden section.
6. If the number of worker processes is different from the default value of 1, then change the number of
worker processes to 1.

Volumes and data sources 27


7. Click OK.
8. Restart IIS Services.

Configuration of SharePoint
When you configure SharePoint, certain privileges are required by user account along with IBM StoredIQ
recommendations. Additionally, SharePoint 2007, 2010, 2013, and 2016 require the configuration of
alternate-access mappings to map IBM StoredIQ requests to the correct websites.
To configure SharePoint, consider these connections and privileges:
Secure Connection
If you want to connect to SharePoint volumes over HTTPS, you can either select the Use SSL
checkbox or add port number 443 after the server name when you set up the volume on IBM
StoredIQ. If you choose the latter option, an example is [Link]. In some cases,
this secure connection can result in some performance degradation due to Secure Socket Layer (SSL)
running large. If you enter the volume information without the 443 suffix, the default connection is
over HTTP.
Note: SharePoint Online connection uses claims-based authentication only. OAuth is not supported
currently.
Privileges
To run policies on SharePoint servers, you must use credentials with Full Control privileges. Use a site
collection administrator to harvest subsites of a site collection.

Privileges required by user account


IBM StoredIQ is typically used with SharePoint for one of these instances: to harvest and treat SharePoint
as a source for policy actions or to use as a destination for policy actions, which means that you can write
content into SharePoint with IBM StoredIQ. Consider these points:
• Attributes are not set or reset on a SharePoint harvest or if you copy from SharePoint.
• Attributes are set only if you copy to SharePoint.
You must denote the following situations:
• If you plan to read only from the SharePoint (harvest and source copies from), then you must use user
credentials with Read privileges on the site and all of the lists and data objects that you expect to
process.
• If you plan to use SharePoint as a destination for policies, you must use user credentials with
Contribute privileges on the site.
• More Privileges for Social Data: If you want to index all the social data for a user profile in SharePoint
2010, then the user credentials must own privileges to Manage Social Data as well.
• Privileges: Use a site collection administrator to ensure that all data is harvested from a site or site
collection.

Alternate-access mappings
Alternate-access mappings map URLs presented by IBM StoredIQ to internal URLs received by Windows
SharePoint Services. An alternate-access mapping is required between the server name and optional port
that is defined in the SharePoint volume definition and the internal URL of the web application. If SSL is
used to access the site, ensure that the alternate-access mapping URL uses https:// as the protocol.
Refer to Microsoft SharePoint 2007, 2010, 2013, or 2016 documentation to configure alternate-access
mappings. These mappings are based on the public URL that is configured by the local SharePoint
administrator and used by the IBM StoredIQ SharePoint volume definitions.
For example, you are accessing a SharePoint volume with the fully qualified domain name, http://
[Link], from the intranet zone. An alternate-access mapping for the public URL
[Link] for the intranet zone must be configured for the SharePoint
2007, 2010, 2013, or 2016 web application that hosts the site to be accessed by the volume definition. If

28 IBM StoredIQ: Data Server Administration Guide


you are accessing the same volume with SSL, the mapping added must be for the URL https://
[Link] instead.
Note: When you configure SharePoint volumes with non-qualified names, you are entering the URL for a
SharePoint site collection or site that is used by IBM StoredIQ in the volume definition. Consider the
following conditions:
• The URL must be valid about the Alternate Access Mappings that are configured in SharePoint.
• If the host name in the URL does not convey the fully qualified domain to authenticate the configured
user, an Active Directory server must be specified. The specified Active Directory must be a fully
qualified domain name and is used for authentication.

Configuring Documentum
Documentum has configuration requirements when it is used as a server platform.
• To run harvests and copy from Documentum servers, you must use the Contributor role.

Installing Documentum client jars to the data server


Documentum can be used as a data source within IBM StoredIQ, but it must be downloaded and then the
RPM installed and run on the target data server.
1. Using this command, create a directory on your target data server to hold the Documentum .JAR files:
mkdir /deepfs/documentum/dfc
2. Copy the Documentum .JAR files to this location on your target data server: /deepfs/
documentum/dfc
3. Download and install rpm-build on your target data server.
a) To locate the rpm-build, go to [Link]
Packages/rpm-build-4.8.0-59.el6.x86_64.rpm
b) Copy rpm-build-4.8.0-59.el6.x86_64.rpm to /deepfs/documentum on your target data
server.
c) Using this command, change directories and install rpm-build: cd /deepfs/documentum
d) Using this command, install rpm-build: rpm --nodeps -i rpm-
build-4.8.0-59.el6.x86_64.rpm
4. Using this command, run the script to create the Documentum package: /usr/local/bin/build-
dfc-client-rpm /deepfs/documentum/dfc
The output is placed in /deepfs/documentum and is named something similar to siq-war-dfc-
client-202.0.0.0p46-1.x86_64.rpm
5. Using this command, run the newly created rpm on your target data server: rpm -i siq-war-dfc-
client-202.0.0.0p46-1.x86_64.rpm
You must build the Documentum RPM only once. Copy the created file siq-war-dfc-
client-202.0.0.0p46-1.x86_64.rpm to each data server you plan to connect to Documentum and
then run the rpm as described in step “5” on page 29.

Configuring NewsGator
When NewsGator is used as a server platform, several privileges must be configured.
Privileges Required by User Account: The user account to harvest or copy from a NewsGator volume must
have the Legal Audit permission on the NewsGator Social Platform Services running on the SharePoint
farm.
1. Log in as an administrator to your SharePoint Central Administration Site.
2. Under Application Management, select Manage Service Applications.
3. In the Manage Service Applications screen, select the NewsGator Social Platform Services row.
4. From the toolbar, select Administrators.

Volumes and data sources 29


5. Add the user account that is used for the NewsGator harvest to the list of administrators. Ensure that
the account has the Legal Audit permission.

Creating primary volumes


A primary volume serves as a primary data source in IBM StoredIQ. You must have at least one primary
volume within your configuration.
1. Go to Administration > Data sources > Specify volumes > Volumes.
2. On the Primary volume list page, click Add primary volumes.
3. Enter the information that is described in the following tables, which are based on your server type.
Individual tables describe the options.
• Chatter primary volumes
• CIFS/SMB or SMB2 (Windows platforms) primary volumes
• Connections primary volumes
• CMIS primary volumes
• Documentum primary volumes
• Domino primary volumes
• Exchange primary volumes
• FileNet primary volumes
• HDFS primary volumes
• IBM Content Manager primary volumes
• Jive primary volumes
• Livelink primary volumes
• NFS V2 and V3 primary volumes
• NewsGator primary volumes
• SharePoint primary volumes
Except for Chatter, Domino®, and Jive volumes, volumes can also be added in IBM StoredIQ
Administrator. However, the set of available configuration options slightly varies. For example, settings
for the synchronization with a governance catalog can be configured only in IBM StoredIQ
Administrator.
Box and Onedrive volumes can be added from IBM StoredIQ Administrator only.
4. Click OK to save the volume.
5. Select one of the following options:
• Add another volume on the same server.
• Add another volume on a different server.
• Finished adding volumes.
This table describes the fields that are available in the Add volume dialog box when you configure
primary volumes.
Note: Case-sensitivity rules for each server type apply. Red asterisks within the user interface denote
the fields.

30 IBM StoredIQ: Data Server Administration Guide


Table 8. CIFS/SMB or SMB2 (Windows platforms) primary volumes
Field Value Notes
Server type Select CIFS (Windows Both SMB and SMB2 are
platform). supported. Depending on the
setup of your SMB server, some
additional SMB configuration
might be required on the IBM
StoredIQ data server. For
details, see the information
about configuring SMB
properties in the Data Server
administration guide.
If you want to preserve
ownership of objects in Copy or
Move actions between CIFS
volumes, you can add an admin
knob as described in the
respective instructions in the
Data Server administration
guide.

Server Enter the fully qualified name of


the server where the volume is
available for mounting.
Connect as Enter the logon ID that is used
to connect and mount the
defined volume.
Password Enter the password that is used
to connect and mount the
defined volume.
Volume Enter the name of the share to
be mounted.
Initial directory Enter the name of the initial With this feature, you can select
directory from which the harvest a volume further down the
must begin. directory tree rather than
selecting an entire volume.
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Validation To validate volume accessibility, When selected (the default


select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.

Volumes and data sources 31


Table 8. CIFS/SMB or SMB2 (Windows platforms) primary volumes (continued)
Field Value Notes
Include directories Specify a regular expression for These directories are defined as
included directories for each sets of "first node" directories,
harvest (if it was specified). relative to the specified (or
implied) starting directory, that
is considered part of the logical
volume.
Start directory Designate a start directory for
the harvest. The start directory
involves volume partitioning to
break up a large volume. If an
initial directory is defined, the
start directory must be
underneath the initial directory.
In the case of directories E-H, E
is the start directory and H is the
end directory.
End directory Determine the end directory for
the harvest. The end directory is
also part of volume partitioning
and is the last directory
harvested.
Access Times Select one of these options:
• Reset access times but do
not synchronize them. (This
setting is the default setting.)
• Do not reset or synchronize
access times.
• Reset and synchronize
access times on incremental
harvests.

32 IBM StoredIQ: Data Server Administration Guide


Table 8. CIFS/SMB or SMB2 (Windows platforms) primary volumes (continued)
Field Value Notes
Constraints Select one of these options:
• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.
• Scope harvests on these
volumes by extension:
Include or exclude data
objects that are based on
extension.

Table 9. NFS v2 and v3 primary volumes


Field Value Notes
Server type Select NFS v2 or NFS v3.
Server Enter the fully qualified name of
the server where the volume is
available for mounting.
Volume Enter the name or names of the
volume to be mounted.
Initial directory Enter the name of the initial With this feature, you can select
directory from which the harvest a volume further down the
must begin. directory tree rather than
selecting an entire volume.
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Volumes and data sources 33


Table 9. NFS v2 and v3 primary volumes (continued)
Field Value Notes
Validation To validate volume accessibility, When selected (the default
select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.
Include directories Specify a regular expression for These directories are defined as
included directories for each sets of "first node" directories,
harvest (if it was specified). relative to the specified (or
implied) starting directory, that
is considered part of the logical
volume.
Start directory Designate a start directory for
the harvest. The start directory
involves volume partitioning to
break up a large volume. If an
initial directory is defined, the
start directory must be
underneath the initial directory.
In the case of directories E-H, E
would be the start directory and
H would be the end directory.
End directory Determine the end directory for
the harvest. The end directory is
also part of volume partitioning
and is the last directory
harvested.
Access times Select one of these options:
• Reset access times but do
not synchronize them. (This
setting is the default setting.)
• Do not reset or synchronize
access times.
• Reset and synchronize
access times on incremental
harvests.

34 IBM StoredIQ: Data Server Administration Guide


Table 9. NFS v2 and v3 primary volumes (continued)
Field Value Notes
Constraints Select one of these options:
• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.
• Scope harvests on these
volumes by extension:
Include or exclude data
objects that are based on
extension.

Table 10. Exchange primary volumes


Field Value Notes
Server type Select Exchange.
Version In the Version list, select the
appropriate version. Options
include 2000/2003, 2007,
2010/2013/2016, and Online.
Server Enter the fully qualified name of If you selected Online as the
the server where the volume is Server option, this field fills
available for mounting. automatically with the online
server name.
For Exchange primary volumes,
it is the fully qualified domain
name where the OWA is.
Multiple Client Access servers
on Exchange 2007 are
supported. The server load must
be balanced at the IP or DNS
level.

Volumes and data sources 35


Table 10. Exchange primary volumes (continued)
Field Value Notes
Mailbox server When you configure multiple For Exchange primary volumes,
client access servers, enter the it is the fully qualified domain
name of one or more mailbox name where the mailbox to be
servers, which are separated by harvested is.
a comma.
If you selected Online as the
Server option, this field is not
available.

Active Directory server Enter the name of the Active It must be a fully qualified
Directory server. Active Directory server.
If you selected Online as the
Server option, this field is not
available.

Protocol To use SSL, select the Protocol If you selected Online as the
checkbox. Server option, the Use SSL
checkbox is automatically
selected, and this field cannot
be edited.
Connect as Enter the logon ID that is used
to connect and mount the
defined volume.
Password Enter the password that is used
to connect and mount the
defined volume.
Volume Enter the name or names of the For Exchange, enter a friendly
volume to be mounted. name for the volume.
Folder Select either of the Mailboxes
or Public folders options.
Initial directory Enter the name of the initial For Exchange, this field must be
directory from which the harvest left blank if you are harvesting
must begin. all mailboxes. If you are
harvesting a single mailbox,
enter the email address for that
mailbox.
Virtual root The name defaults to the
correct endpoint for the
selected Exchange version.
Personal archives Select Harvest personal This checkbox is available only
archive to harvest personal when Exchange
archives. 2010/2013/2016 or Online is
selected.

36 IBM StoredIQ: Data Server Administration Guide


Table 10. Exchange primary volumes (continued)
Field Value Notes
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Remove journal envelope When selected, the journal


envelope is removed.
Validation To validate volume accessibility, When selected (the default
select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.
Include directories Specify a regular expression for These directories are defined as
included directories for each sets of "first node" directories,
harvest (if it was specified). relative to the specified (or
implied) starting directory that
is considered part of the logical
volume.
Start directory Designate a start directory for The parameters are date ranges
the harvest. The start directory that are used to scope the
involves volume partitioning to harvest, the format of which is
break up a large volume. If an YYYY-MM-DD.
initial directory is defined, the
start directory must be
underneath the initial directory.
In the case of directories E-H, E
would be the start directory and
H would be the end directory.
End directory Determine the end directory for The parameters are date ranges
the harvest. The end directory is that are used to scope the
also part of volume partitioning harvest, the format of which is
and is the last directory YYYY-MM-DD.
harvested.

Volumes and data sources 37


Table 10. Exchange primary volumes (continued)
Field Value Notes
Constraints Select one of these options:
• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.

Table 11. SharePoint primary volumes.


Prerequisites: For SharePoint volume prerequisites and configuration information, see
“Configuration of SharePoint” on page 28 and “Special note about SharePoint volumes” on page 65.

Field Value Notes


Server type SharePoint Required.
Version Select one of these servers: Required.
2003, 2007, 2010, 2013, 2016,
or Online.
Server The fully qualified name of the Required. When you add
SharePoint server. SharePoint volumes that contain
spaces in the URL, see Special
Note: Adding SharePoint
Volumes.
Active Directory server The name of the Active Optional. Specify the fully
Directory server. qualified Active Directory server
name. This option is not
available for SharePoint Online.

38 IBM StoredIQ: Data Server Administration Guide


Table 11. SharePoint primary volumes.
Prerequisites: For SharePoint volume prerequisites and configuration information, see
“Configuration of SharePoint” on page 28 and “Special note about SharePoint volumes” on page 65.
(continued)
Field Value Notes
Protocol Select Use SSL only if SSL is Optional.
enabled for this SharePoint
If SSL is enabled on the
server. For SharePoint Online,
SharePoint server and you do
this option is automatically
not select this option, no volume
selected and cannot be edited.
is created and the HTTP status
code 301 Moved
Permanently is returned. To
fix the issue, select the option.
If SSL is not enabled on the
SharePoint server and you
select this option, no volume is
created and the socket error
[Errno 111] Connection
Refused is returned. To fix this
issue, clear the Use SSL
checkbox.

Connect as Enter the name of a user with Required. Use a site collection
the required permissions for administrator account.
that site collections. Use the
No volume can be added if the
following syntax:
validation of the credentials
• SharePoint Online: fails, which can happen for the
following reasons:
userid@Microsoft_cloudname
.com • The user does not exist or
does not have the required
• Other SharePoint versions: permissions.
Active Directory Domain • The password is not correct.
Name\username
The HTTP status code is usually
401 Unauthorized. However,
for SharePoint Online, the HTTP
status code 400 Bad Request
is returned for insufficient
permissions.

Password Enter the password for the user Required.


specified in Connect as.
Volume Enter the URL of the SharePoint Required. Do not include the
site collection, for example: / SharePoint server name in the
portal/site URL, otherwise the URL cannot
be located on the server and
thus no volume is created.
When you add SharePoint
volumes that contain spaces in
the URL, see Special Note:
Adding SharePoint Volumes.

Volumes and data sources 39


Table 11. SharePoint primary volumes.
Prerequisites: For SharePoint volume prerequisites and configuration information, see
“Configuration of SharePoint” on page 28 and “Special note about SharePoint volumes” on page 65.
(continued)
Field Value Notes
Initial directory Enter the name of the subsite Optional.
from which you want the harvest
• With this feature, you can
to start.
select a volume further down
the directory tree rather than
selecting an entire volume.
• When you add SharePoint
volumes that contain spaces
in the URL, see Special Note:
Adding SharePoint Volumes.

Index options Select either or both of the Both options are selected by
Index options checkboxes. default.
• Include system metadata for Tip: Leave the Include
data objects within metadata for contained
containers objects checkbox selected to
• Include content tagging and have metadata for objects in
full-text index containers added to the
metadata index. To avoid
creating a full-text index for the
entire volume, clear the Include
content tagging and full-text
index checkbox. Create a full-
text index for a subset of data
later by running a Step-up
Analytics action.
For SharePoint Online, full-text
indexing of OneNote notebook
objects, that is, Notes, is not
supported currently. FSMD-
based searches for these files
are supported.

Subsites To check all sites and subsites Optional.


of the site collection for data
objects, select Recurse into
subsites.
Versions To harvest all document Optional. IBM StoredIQ
versions, select Include all supports indexing versions from
versions. SharePoint. For more
information, see Special Note:
Adding SharePoint Volumes.
Validation To validate volume accessibility, When selected (the default
select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.

40 IBM StoredIQ: Data Server Administration Guide


Table 11. SharePoint primary volumes.
Prerequisites: For SharePoint volume prerequisites and configuration information, see
“Configuration of SharePoint” on page 28 and “Special note about SharePoint volumes” on page 65.
(continued)
Field Value Notes
Include directories Specify a regular expression for These directories are defined as
included directories for each sets of "first node" directories,
harvest (if it was specified). relative to the specified (or
implied) starting directory, that
is considered part of the logical
volume.
Start directory Designate a start directory for
the harvest. The start directory
involves volume partitioning to
break up a large volume. If an
initial directory is defined, the
start directory must be
underneath the initial directory.
In the case of directories E-H, E
would be the start directory and
H would be the end directory.
End directory Determine the end directory for
the harvest. The end directory is
also part of volume partitioning
and is the last directory
harvested.
Constraints Select one of these options:
• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.

Volumes and data sources 41


Table 12. Documentum primary volumes.
Prerequisites: Before you can add Documentum volumes, you must add the Documentum server.
For more information, see tsk_addingadocumentumserverasadatasource.dita.

Field Value Notes


Server type Select Documentum. For Documentum, you must
specify the doc broker. See
“Configuring Documentum” on
page 29.
Doc base Enter the name of the A Documentum repository
Documentum repository. contains cabinets, and cabinets
contain folders and documents.
Connect as Enter the logon ID that is used
to connect and mount the
defined volume.
Password Enter the password that is used
to connect and mount the
defined volume.
Volume Enter the name or names of the For Documentum, enter a
volume to be mounted. friendly name for the volume.

Harvest To enable harvesting all Important: If you do not select


document versions, select this option for the initial harvest,
Harvest all document versions. changing the setting later does
not have an effect when the
volume is reharvested. As a
workaround, create a new
volume and ensure that the
Harvest all document versions
option is set before you start
harvesting.

Initial directory Enter the name of the initial With this feature, you can select
directory from which the harvest a volume further down the
must begin. directory tree rather than
selecting an entire volume.
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Validation To validate volume accessibility, When selected (the default


select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.

42 IBM StoredIQ: Data Server Administration Guide


Table 12. Documentum primary volumes.
Prerequisites: Before you can add Documentum volumes, you must add the Documentum server.
For more information, see tsk_addingadocumentumserverasadatasource.dita.
(continued)
Field Value Notes
Include directories Specify a regular expression for These directories are defined as
included directories for each sets of "first node" directories,
harvest (if it was specified). relative to the specified (or
implied) starting directory that
is considered part of the logical
volume.
Constraints Select one of these options:
• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.

Table 13. Domino primary volumes


Field Value Notes
Server type Select Domino. For Domino, you must first
upload at least one [Link]. See
Adding Domino as a Primary
Volume.

Volumes and data sources 43


Table 13. Domino primary volumes (continued)
Field Value Notes
Server Enter the fully qualified name of For Domino, select the
the server where the volume is appropriate user name, which
available for mounting. was entered with the
Configuration subtab in the
It can happen that the data
Lotus Notes® user
server cannot find the Domino
administration area.
server based on this information
because the DNS resolution
fails. In this case, ping the
Domino server. If the ping is
successful, add an entry to the
data server's /etc/hosts file.
The entry must consist of the
common name portion of the
Domino server name (omit the
domain portion) and the IP
address associated with the
Domino server, for example:

[Link] NALLN999

Connect as Enter the logon ID that is used For Domino, select the user
to connect and mount the name for the primary user ID.
defined volume. The user ID must be configured
on the System Configuration
screen under the Lotus Notes
user administration link.
Password Enter the password that is used For Domino, enter the password
to connect and mount the for the primary user ID.
defined volume.
Volume Enter the name or names of the For Domino, enter a friendly
volume to be mounted. name for the volume.
Harvest • To harvest mailboxes, select This option obtains the list of all
the Harvest mailboxes known Domino users and their
option. NSFs. It then harvests those
mailboxes unless it was pointed
• To harvest mail journals, to a single mailbox with the
select the Harvest mail initial directory.
journals option.
• To harvest all applications,
select the Harvest all
applications option.

Initial directory Enter the name of the initial


directory.

44 IBM StoredIQ: Data Server Administration Guide


Table 13. Domino primary volumes (continued)
Field Value Notes
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Validation To validate volume accessibility, When selected (the default


select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.

Include directories Specify a regular expression for These directories are defined as
included directories for each sets of "first node" directories,
harvest (if it was specified). relative to the specified (or
implied) starting directory, that
is considered part of the logical
volume.
Start directory Designate a start directory for
the harvest. The start directory
involves volume partitioning to
break up a large volume. If an
initial directory is defined, the
start directory must be
underneath the initial directory.
In the case of directories E-H, E
would be the start directory and
H would be the end directory.
End directory Determine the end directory for
the harvest. The end directory is
also part of volume partitioning
and is the last directory
harvested.

Volumes and data sources 45


Table 13. Domino primary volumes (continued)
Field Value Notes
Constraints Select one of these options:
• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.

Table 14. FileNet primary volumes


Field Value Notes
Server type Select FileNet. Within IBM StoredIQ Data
Server, the FileNet® domain
must be configured before any
FileNet volumes are created.
See “Configuring FileNet” on
page 66.
FileNet config Select the FileNet server you For more information, see
would like to use for this “Configuring FileNet” on page
configuration. 66.
Connect as Enter the logon ID that is used
to connect and mount the
defined volume.
Password Enter the password that is used
to connect and mount the
defined volume.
Domain Domain name automatically
populates.
Object store Select an object store. The object store must exist
before you create a FileNet
primary volume.
Volume Enter the name or names of the
volume to be mounted.

46 IBM StoredIQ: Data Server Administration Guide


Table 14. FileNet primary volumes (continued)
Field Value Notes
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Constraints Select one of these options:


• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.

Table 15. NewsGator primary volumes


Field Value Notes
Server type Select NewsGator.
Server Enter the fully qualified name of
the server where the volume is
available for mounting.
Protocol To use SSL, select the Protocol
checkbox.
Connect as Enter the logon ID that is used
to connect and mount the
defined volume.
Password Enter the password that is used
to connect and mount the
defined volume.
Volume Enter the name or names of the Enter a friendly name for the
volume to be mounted. volume.

Volumes and data sources 47


Table 15. NewsGator primary volumes (continued)
Field Value Notes
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Validation To validate volume accessibility, When selected (the default


select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.

Table 16. Livelink primary volumes.


Prerequisite: A copy of the [Link] file from a Livelink API installation must be available in
the /usr/local/IBM/ICI/vendor directory on each data server in your deployment. Usually, you
can find this file on the Livelink server in the C:\OPENTEXT\application\WEB-INF\lib
directory. However, the path might be different in your Livelink installation.

Field Value Notes


Server type Select Livelink.
Server Enter the fully qualified name of
the server where the volume is
available for mounting.
Port Enter the port number to be
used.
Database Enter the name of the database.
Search slice Enter the name of the search
slice.
Connect as Enter the logon ID that is used
to connect and mount the
defined volume.
Password Enter the password that is used
to connect and mount the
defined volume.
Volume Enter the name or names of the
volume to be mounted.
Initial directory Enter the search slice and the With this feature, you can select
name of the initial directory a volume further down the
from which the harvest must directory tree rather than
begin. selecting an entire volume.

48 IBM StoredIQ: Data Server Administration Guide


Table 16. Livelink primary volumes.
Prerequisite: A copy of the [Link] file from a Livelink API installation must be available in
the /usr/local/IBM/ICI/vendor directory on each data server in your deployment. Usually, you
can find this file on the Livelink server in the C:\OPENTEXT\application\WEB-INF\lib
directory. However, the path might be different in your Livelink installation.
(continued)
Field Value Notes
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Validation To validate volume accessibility, When selected (the default


select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.
Include directories Specify a regular expression for These directories are defined as
included directories for each sets of "first node" directories,
harvest (if it was specified). relative to the specified (or
implied) starting directory that
is considered part of the logical
volume.
Start directory Designate a start directory for
the harvest. The start directory
involves volume partitioning to
break up a large volume. If an
initial directory is defined, the
start directory must be
underneath the initial directory.
In the case of directories E-H, E
would be the start directory and
H would be the end directory.
End directory Determine the end directory for
the harvest. The end directory is
also part of volume partitioning
and is the last directory
harvested.

Volumes and data sources 49


Table 16. Livelink primary volumes.
Prerequisite: A copy of the [Link] file from a Livelink API installation must be available in
the /usr/local/IBM/ICI/vendor directory on each data server in your deployment. Usually, you
can find this file on the Livelink server in the C:\OPENTEXT\application\WEB-INF\lib
directory. However, the path might be different in your Livelink installation.
(continued)
Field Value Notes
Constraints Select one of these options:
• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.
• Scope harvests on these
volumes by extension:
Designate which data objects
must be harvested by entering
those object extensions.

Table 17. Jive primary volumes


Field Value Notes
Server type Select Jive.
Server Enter the fully qualified name of
the server where the volume is
available for mounting.
Protocol To use SSL, select the Protocol
checkbox.
Connect as Enter the logon ID that is used
to connect and mount the
defined volume.
Password Enter the password that is used
to connect and mount the
defined volume.
Volume Enter the name or names of the
volume to be mounted.

50 IBM StoredIQ: Data Server Administration Guide


Table 17. Jive primary volumes (continued)
Field Value Notes
Initial directory Enter the name of the initial With this feature, you can select
directory from which the harvest a volume further down the
must begin. directory tree rather than
selecting an entire volume.
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Versions To harvest all document


versions, select Include all
versions.
Validation To validate volume accessibility, When selected (the default
select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.
Include directories Specify a regular expression for These directories are defined as
included directories for each sets of "first node" directories,
harvest (if it was specified). relative to the specified (or
implied) starting directory that
is considered part of the logical
volume.
Start directory Designate a start directory for
the harvest. The start directory
involves volume partitioning to
break up a large volume. If an
initial directory is defined, the
start directory must be
underneath the initial directory.
In the case of directories E-H, E
would be the start directory and
H would be the end directory.
End directory Determine the end directory for
the harvest. The end directory is
also part of volume partitioning
and is the last directory
harvested.

Volumes and data sources 51


Table 17. Jive primary volumes (continued)
Field Value Notes
Constraints Select one of these options:
• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.

Table 18. Chatter primary volumes


Field Value Notes
Server type Select Chatter. For Chatter, see “Configuring
Chatter messages” on page
66.
Connect as Enter the logon ID that is used
to connect and mount the
defined volume.
Password Enter the password that is used
to connect and mount the
defined volume.
Auth token Enter the token that is used to The auth token must match the
authenticate the Chatter user name that is used in the
volume. Connect as field. Auth tokens
can be generated online on
Salesforce. See Configuring
chatter messages.
Volume Enter the name or names of the
volume to be mounted.
Initial directory Enter the name of the initial With this feature, you can select
directory from which the harvest a volume further down the
must begin. directory tree rather than
selecting an entire volume.

52 IBM StoredIQ: Data Server Administration Guide


Table 18. Chatter primary volumes (continued)
Field Value Notes
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Validation To validate volume accessibility, When selected (the default


select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.
Include directories Specify a regular expression for These directories are defined as
included directories for each sets of "first node" directories,
harvest (if it was specified). relative to the specified (or
implied) starting directory, that
is considered part of the logical
volume.
Start directory Designate a start directory for
the harvest. The start directory
involves volume partitioning to
break up a large volume. If an
initial directory is defined, the
start directory must be
underneath the initial directory.
In the case of directories E-H, E
would be the start directory and
H would be the end directory.
End directory Determine the end directory for
the harvest. The end directory is
also part of volume partitioning
and is the last directory
harvested.

Volumes and data sources 53


Table 18. Chatter primary volumes (continued)
Field Value Notes
Constraints Select one of these options:
• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.

Table 19. IBM Content Manager primary volumes


Field Value Notes
Server type Select IBM Content Manager.
Server Enter the fully qualified host
name of the library server
database.
Port Enter the port that is used to
access the library server
database.
Repository Enter the name of the library
server database.
Database type Select the type of database that
is associated with the volume.
Options include DB2 and
Oracle. By default, DB2 is
selected.
Schema Enter the schema of the library
server database.
Remote database Enter the name of the remote Optional.
database.
Connection String Enter any additional Optional.
parameters.
Harvest itemtype Enter the name of the item Required.
types to be harvested,
separated by commas.

54 IBM StoredIQ: Data Server Administration Guide


Table 19. IBM Content Manager primary volumes (continued)
Field Value Notes
Copy to itemtype Only SiqDocument is For more information, see “IBM
supported. Content Manager attributes” on
page 69
Connect as Enter the logon ID that is used
to connect and mount the
defined volume.
Note: This Content Manager
user ID must have access to all
documents to be able to create
documents in the
SiqDocument item type for
copy to. If the SiqDocument
item type does not exist, a
Content Manager administration
ID must be used, as this ID
creates the item type when the
volume is created.

Password Enter the password that is used


to connect and mount the
defined volume.
Volume Enter the name or names of the
volume to be mounted.
Index options Select either or both of the Both options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Validation To validate volume accessibility, When selected (the default


select Validation. state), IBM StoredIQ tests to
see whether the volume can be
accessed.

Volumes and data sources 55


Table 19. IBM Content Manager primary volumes (continued)
Field Value Notes
Constraints Select one of these options:
• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.
• Scope harvests on these
volumes by extension:
Designate which data objects
must be harvested by entering
those object extensions.

Table 20. CMIS primary volumes


Field Value Notes
Server type Select CMIS.
Server Enter the fully qualified name of
the server where the volume is
available for mounting.
Port Enter the name of the port.
Repository Enter the name of the
repository.
Service In the Service text box, enter
the name of the service.
Protocol To use SSL, select the Protocol
checkbox.
Connect as Enter the logon ID that is used
to connect and mount the
defined volume.
Password Enter the password that is used
to connect and mount the
defined volume.

56 IBM StoredIQ: Data Server Administration Guide


Table 20. CMIS primary volumes (continued)
Field Value Notes
Volume Enter the name or names of the
volume to be mounted.
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Validation To validate volume accessibility, When selected (the default


select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.

Constraints Select one of these options:


• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.

Table 21. HDFS primary volumes


Field Value Notes
Server type Select HDFS. Required.
Server Enter the host name or IP Required.
address.
Port Enter the port number.
Repository Enter the name of the
repository.

Volumes and data sources 57


Table 21. HDFS primary volumes (continued)
Field Value Notes
Option string This option is supported: This option is used to indicate
VerifyCertificate=True. that the validity of the HDFS
server's SSL certificate is
verified when SSL is used.
Values are True, False, or
default value. If no value is
specified, value is False. To
validate the certificate on the
HDFS server, the user needs to
specify this option and set the
value to True.
Protocol To use SSL, select the Protocol
checkbox.
Connect as Enter the logon ID that is used
to connect and mount the
defined volume.
Password Enter the password that is used Authentication to HDFS is not
to connect and mount the supported. If your HDFS server
defined volume. requires a password, StoredIQ
is not able to connect to it.
Volume Enter the name or names of the
volume to be mounted.
Initial directory Enter the name of the initial To avoid the interrogator
directory from which the harvest timeout issue, see Note at the
must begin. end of this table.
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Validation To validate volume accessibility,


select this option.
Include directories
Start directory Designate a start directory for
the harvest. The start directory
involves volume partitioning to
break up a large volume. If an
initial directory is defined, the
start directory must be
underneath the initial directory.
In the case of directories E-H, E
would be the start directory and
H would be the end directory.

58 IBM StoredIQ: Data Server Administration Guide


Table 21. HDFS primary volumes (continued)
Field Value Notes
End directory Determine the end directory for
the harvest. The end directory is
also part of volume partitioning
and is the last directory
harvested.
Contraints Select one of these options:
• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.
• Scope harvests on these
volumes by extension:
Designate which data objects
must be harvested by entering
those object extensions.

Note: If you harvest HDFS volumes with many files in a directory, then an interrogator timeout might
occur resulting in a Skipped directory exception in the harvest audit. HDFS responds to StoredIQ
slowly when it handles large directories and processes more responses from HDFS. The slow response
from HDFS is caused by high CPU usage on HDFS NameNode. Therefore, if interrogator timeout occurs
and high CPU usage on the HDFS server is observed, you can allocate more CPU resources to the HDFS
server.
To avoid the interrogator timeout issue, you can also limit the file number to 250,000 files in a
directory. Since each directory has its own timeout, having fewer files in a single directory ensures
efficient operation. Splitting large directories into many small ones also helps resolve the interrogator
timeout issues. For example, 1,000,000 files that are equally distributed into 10 directories have
fewer risks of timeouts than if they are in one directory.

Table 22. Connections primary volumes


Field Value Notes
Server Type Select Connections in the Required
Server type list.

Volumes and data sources 59


Table 22. Connections primary volumes (continued)
Field Value Notes
Server Enter the fully qualified domain Required
name of the server from which
the volume is available for
mounting.
Class name Enter Required

[Link].
[Link].
ibmconnectionsconn.
IBMConnections

Repository name Enter Required

[Link].
[Link].
ibmconnectionsconn

Option string
Connect as Enter the user name of the Required
account that is set up with
admin and search-admin
privileges on the Connections
server.
Password Enter the password of the Required
account that is set up with
admin and search-admin
privileges on the Connections
server.
Volume Enter any name. Required
Initial Directory
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Volume Enter the name or names of the


volume to be mounted.
Validation To validate volume accessibility, When selected (the default
select this option. state), IBM StoredIQ tests to
see whether the volume can be
accessed.
Include directories Specify a regular expression for These directories are defined as
included directories for each sets of "first node" directories,
harvest (if it was specified). relative to the specified (or
implied) starting directory that
is considered part of the logical
volume.

60 IBM StoredIQ: Data Server Administration Guide


Table 22. Connections primary volumes (continued)
Field Value Notes
Start directory Designate a start directory for
the harvest. The start directory
involves volume partitioning to
break up a large volume. If an
initial directory is defined, the
start directory must be
underneath the initial directory.
In the case of directories E-H, E
is the start directory and H is the
end directory.
End directory Determine the end directory for
the harvest. The end directory is
also part of volume partitioning
and is the last directory
harvested.
Access Times Select one of these options:
• Reset access times but do
not synchronize them. (This
setting is the default setting.)
• Do not reset or synchronize
access times.
• Reset and synchronize
access times on incremental
harvests.

Constraints Select one of these options:


• Only use __ connection
process (es): Specify a limit
for the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.
• Scope harvests on these
volumes by extension:
Include or exclude data
objects that are based on
extension.

Volumes and data sources 61


Configuring SMB properties
Depending on the configuration of your SMB server, you might need to change the SMB settings in IBM
StoredIQ Data Server to make the settings match.
To change the SMB settings on the data server to which you want to add the volume:
1. Using an SSH tool, log in to the data server as root.
2. Create a [Link] file in the /usr/local/siqsmb folder by copying the original
properties file. Use this command:

cp /usr/local/siqsmb/lib/[Link]
/usr/local/siqsmb/[Link]

The properties file contains a set of SMB properties that you can use to adjust your SMB configuration.
For information about additional properties for further configuration, contact IBM Support.
[Link]
Determines whether client signing for interprocess communication (IPC) connections is enforced.
The default value is true.
This means that, although the data server is configured with signing not being required and not
being supported, IPC connections to the SMB server still require signing by default. If the SMB
servers does not support signing for IPC connections, the data server cannot connect to that server
during volume creation unless you set this property to false.
Changed setting example: [Link]=false
[Link]
Determines the maximum number of directories and files to be returned with each request of the
TRANS2_FIND_FIRST/NEXT2 operation. The default value is 200.
Depending on the characteristics of the files and directories on the SMB server, you might want to
adjust this value for performance reasons. For example, on higher latency networks a lower value
can result in better performance.
Changed setting example: [Link]=300
[Link]
Enables SMB signing if available. The default value is false.
If the SMB server requires SMB signing, you must set this property to true to have the data server
as a JCIFS client negotiate SMB signing with that SMB server. If the SMB server does not require
SMB signing but supports it, you can set this property to true for signing to occur. Otherwise, SMB
signing is disabled.
Changed setting example: [Link]=true
[Link]
Determines whether client signing in general is enforced. The default value is false.
If the SMB server does not require and does not support signing, setting this property to true
causes the connection to the SMB server to fail. Only set this property to true if one of these
security policies is set on the SMB server:
• Microsoft network server: Digitally sign communication (always)
• Microsoft network server: Digitally sign communication (if client agrees)
Changed setting example: [Link]=true
[Link].enableSMB2 (deprecated)
Enables SMB2 support. The default value is true.
Changed setting example: [Link].enableSMB2=false
This setting is deprecated as of IBM StoredIQ [Link].

62 IBM StoredIQ: Data Server Administration Guide


[Link].disableSMB1 (deprecated)
Disables SMB1 support. The default value is false.
Changed setting example: [Link].disableSMB1=true
This setting is deprecated as of IBM StoredIQ [Link].
[Link]
Determines the minimum protocol version that is to be used. The default value is SMB1 if this
property is not set. Possible values are SMB1, SMB202, or SMB210.
Changed setting example: [Link]=SMB202
[Link]
Determines the maximum protocol version that is to be used. The default value is SMB210 if this
property is not set. Possible values are SMB1, SMB202, or SMB210.
Changed setting example: [Link]=SMB210
[Link]
Disables Distributed File System (DFS) referrals. The default value is false.
In non-domain environments, you might want to set this property to true to disable domain-based
DFS referrals. Domain-based DFS referrals normally run when the data server as a JCIFS client
first tries to resolve a path. In non-domain environments, these referrals time out causing a long
startup delay.
Changed setting example: [Link]=true
[Link]
Determines the connection timeout, that is the time period in milliseconds that the client waits to
connect to a server. The default value is 35000.
Changed setting example: [Link]=70000
[Link]
Determines the socket timeout, that is the time period in milliseconds after which sockets are
closed if there is no activity. The default value is 35000.
Changed setting example: [Link] =70000
[Link]
Determines the timeout for SBM responses, that is the time period in milliseconds that the client
waits for the server to respond to a request. The default value is 30000.
Changed setting example: [Link]=60000
[Link]
Determines the timeout for SMB sessions, that is the time period in milliseconds after which the
session is closed if there is not activity. The default value is 35000.
Changed setting example: [Link]=70000
3. Edit the /usr/local/siqsmb/[Link] file.
4. Locate the property that you want to change, uncomment it, and set the appropriate value.
5. Restart services using this command: service deepfiler restart
6. Exit the data server.

Adding an SMB1 server as primary volume


In IBM StoredIQ, SMB1 and SMB2 are enabled by default, with the first choice being SMB2 connections.
If the CIFS server that you add as a primary volume does not support SMB2, SMB1 is used. However, if
the server supporting SMB1 only does also not provide client signing, you must disable IPC client signing

Volumes and data sources 63


on the data server on which the respective CIFS volume is defined by setting the
[Link] property to false:

[Link]=false

Enabling ownership preservation for objects on CIFS volumes


To preserve the ownership of objects in Copy and Move actions between CIFS volumes, add an admin
knob.
The credentials used for connecting to the target server must be defined in the server's Administrators
group.
Without preserving ownership, the copied or moved object on the target server is owned by the user
whose credentials are used for connecting to the target server. With ownership preservation enabled,
ownership is defined as follows:
• Objects are copied or moved between CIFS servers in different domains but the user name of the
source object owner is defined in both domains. In this case, the user in the target server domain owns
the copied or moved object. For example, the user JohnDoe is defined in the source domain Support
and in the target domain Service. Thus, objects on the server in the source domain are owned by user
Support/JohnDoe. After being copied or moved to the server in the target domain, the objects are
owned by user Service/JohnDoe.
• Objects are copied or moved between CIFS servers in different domains but the user name of the
source object owner is not defined in the target domain. In this case, the copied or moved object on the
target server is owned by the user whose credentials are used for connecting to the target server. For
example, the source objects are owned by the local user JohnDoe. This user is not defined in the target
domain. The credentials of user JaneDoe are used for connecting to the target server. After being
copied or moved to the server in the target domain, the objects are owned by user JaneDoe.
• If the owner information on the target server cannot be resolved, for example, because a slow
connection caused a query timeout, the object on the target server is owned by the target's Domain
Administrator.
1. Using an SSH tool, log in to the data server VM as root.
2. To insert the admin knob, open a shell and run the following command:

psql -U dfuser dfdata -c "INSERT INTO [Link]


(name,value,description,valuetype,use) VALUES ('smb_set_user_policy', 'smart', 'user policy
copy/move', 'str', 2);"

3. Restart services by running this command:

service deepfiler restart

At any time, you can disable this capability by removing the admin knob with the following command:

psql -U dfuser dfdata -c "DELETE FROM [Link] where name='smb_set_user_policy';"

This also requires restarting services.

Configuring Exchange 2007 Client Access Server support


The system supports the harvest of multiple Client Access Servers (CAS) when you configure Exchange
2007 primary volumes. This feature does not support redirection to other CAS/Exchange clusters or
autodiscovery protocol.
1. Go to Administration > Data sources > Volumes > Primary > Add primary volumes.
2. In the Server type list, select Exchange.
3. In the Version list, select 2007.
4. In the Server text box, type the name of the Exchange server. This server must be load-balanced at
the IP or DNS level.

64 IBM StoredIQ: Data Server Administration Guide


5. In the Mailbox server: text box, enter the name of one or more mailbox servers, which are separated
by a comma and a space.
6. Complete the remaining fields for the primary volume, and then click OK.

Adding Domino as a primary volume


Domino volumes can be added as primary volumes.
1. Add a Lotus Notes user by uploading its user ID file in Lotus Notes User Administration on the
Administration > Configuration tab.
• If you want to harvest a user’s mailbox, add the user ID file for that user.
• If you want to harvest multiple mailboxes within one volume definition, add the administrator’s ID
file.
• If the mailboxes have encrypted emails or NSFs, then you need each user’s user ID file to decrypt a
user’s data.
2. Point the volume to the Domino server. If a single mailbox must be harvested, set the initial directory
to be the path to the mailbox on the Domino server, such as mail\USERNAME.
3. To harvest mailboxes, select the Harvest mailboxes option, which obtains the list of all known Domino
users and their NSFs. It then harvests those mailboxes unless it was pointed to a single mailbox by
using the initial directory.
4. To harvest all mail journals, select the Harvest mail journals option.
5. To harvest all mail applications, select the Harvest all applications option, which looks at all NSFs,
including mail journals, on the Domino server.

Special note about SharePoint volumes


Certain fields must be configured when SharePoint volumes are added.
IBM StoredIQ supports the entire sites portion of a Sharepoint URL for the volume /sites/main_site/
sub_site in the Volume field when you add a SharePoint volume. However, if the SharePoint volume
URL contains spaces, then you must also use the Server, Volume, and Initial directory fields in the Add
volume dialog box in addition to the required fields Server type, Server, Connect as, and Password. For
example, the SharePoint volume with the URL [Link]
autoteamsite1/Attribute Harvest WikiPages Library/ would require the fields in the
following table because of the spaces in the URL.

Table 23. SharePoint volumes as primary volumes: Fields and examples


Primary volume field Example
Server [Link]
Volume /sitestest/autoteamsite1
Initial directory Attribute Harvest Wiki Pages Library

Performance conditions for using versions


When you add a primary volume, you define the volume by setting certain properties. If a SharePoint
volume is added, you have the option of indexing different versions of data objects on that volume.
Since most versions of any object share full-text content and attributes, the effort in processing them and
maintaining an updated context for the version history of an object in the index is duplicated. Additionally,
if you enable version feature on a SharePoint volume, the API itself causes extra overhead in fetching data
and metadata for older versions.
• For each object, an extra API call must be made to get a list of all its versions.
• To fetch attributes for the older versions of an object, an API call must be made for each attribute that
needs to be indexed.

Volumes and data sources 65


Limitations of SharePoint volumes
For SharePoint volumes, these limitations and specific warnings must be carefully considered.
• If data that is essential to the functioning of the SharePoint site as an application is deleted, the site
might become unusable. For example, if you delete stylesheets and forms, you might get errors when
you try to display certain pages on the site.
• It is possible to delete objects that are normally not visible. For example, documents that are filtered by
views might not be visible within SharePoint. Regardless of their visibility, these data objects are
indexed during harvests and can appear in infosets when responsive.
• The delete action is supported only for files that are held in document libraries. Other SharePoint object
types can be present in an infoset, and they are audited as an unsupported operation. No folders, sites,
or document libraries are deleted.
• Delete occurs at the system-file level. It means that the deletion of an older version or of a contained
object results in deletion of all versions and the containing file and all of the objects it contains. The
volume index is updated to reflect this change.
• The delete action can prevent the deletion of items that were accessed or modified since the previous
harvest. For SharePoint volumes, this option is ignored. Recently accessed or modified files are deleted.
• It is not possible to delete files that are currently checked out.

Configuring FileNet
By providing the configuration values for a FileNet domain, you are supplying the values that are needed
to bootstrap into a domain.
Within IBM StoredIQ Data Server, the FileNet domain must be configured prior to any FileNet volumes
being created.
Note: With regards to object storage for FileNet cluster set ups, use other storage mechanisms than
database (BLOB) storage. Additionally, files larger than 100MB should not be written to database storage.
1. Go to Administration > Data sources > Specify Servers > FileNet domain configurations.
2. Click Add new FileNet domain configuration
The FileNet domain configuration editor page appears.
3. In the FileNet domain configuration editor page, configure these fields:
a) In the Configuration name text box, enter the configuration name for this server.
b) In the Server name text box, enter the server name.
c) In the Connection list, select the connection type.
d) In the Port text box, enter the port number.
e) In the Path text box, enter the path for this server.
f) In the Stanza text box, enter the stanza information for this server.
4. Click OK to save your changes.

Configuring Chatter messages


Within Chatter, the default administrator profile does not have the Manage Chatter Messages
permission, but the appropriate permissions are required to harvest private messages.
A user must have certain administrative permissions when that user account is used in the Connect as
text box in Chatter. When you set up a Chatter user account to harvest and run actions against Chatter,
you must use an account with the built-in system administrator profile. In general, however, these
administrative permissions must be assigned to the account you use:
• API enabled
• Manager Chatter Messages (required if you want to harvest Chatter Private Messages)
• Manage Users

66 IBM StoredIQ: Data Server Administration Guide


• Moderate Chatter
• View All Data
• For Chatter administrators who use the Auth token option, see how to set up a sandbox account.

Adding a Documentum server as a data source


A Documentum server can be added as a data source and used as any other primary volume.
1. Using an SSH tool, turn on the Documentum license with these commands:
a) psql -U dfuser -d dfdata -c "update productlicensing set pl_isactive =
true where pl_product = 'documentum';"
b) service deepfiler restart
2. In a browser, log in to the IBM StoredIQ data server.
3. Resolve the server name. Click Administration > Configuration > Network settings > Server name
resolution.
a) Enter the Doc broker settings. In the Host area, enter the Documentum host name, such as
[Link]
If you have more than one host, enter a single host per line. IP addresses can also be used.
b) Click OK.
c) Using an SSH tool, connect to your target data server and edit /etc/hosts. Type ip of
dataserver hostname entered in step 3.a
Use one entry line per host name.
4. Restart services by using either of these methods:
a) Click Administration > Dashboard > Controller. Scroll to the bottom of the page and click Restart
services.
b) Using an SSH tool, enter service deepfiler restart
5. Add Documentum as a primary volume.
You can do this in IBM StoredIQ Data Server or in IBM StoredIQ Administrator. For more information
about adding the volume in IBM StoredIQ Data Server, see “Creating primary volumes” on page 30.
For more information about adding the volume in IBM StoredIQ Administrator, see the topic about
adding primary volumes in the IBM StoredIQ Administrator documentation.

Configuration of IBM Connections


IBM Connections can be harvested and the Copy from action to a CIFS target is supported. Discovery
Exports are also supported.
Note: Not all Profile fields are harvested, such as mobile number, pager number, and fax number. Custom
attributes are supported. Libraries in Connections are links to FileNet objects; these files can be
harvested.
The Copy from action is supported only with a CIFS target. Any harvested Connections instance has the
following directory structure. It is a logical structure of hierarchy, not the actual way that data is stored.

Home
Communities
Files
Forums
Wikis
Activities
Blogs
Status
Bookmarks
Events
Comments
Profiles

Note: When you create a Connections volume, the use of an initial directory, Start directory or End
directory beyond two levels of recursion, is not supported. For example, Home/Files is supported, but

Volumes and data sources 67


Home/Files/User1 is not. Additionally, harvest scoping, which is the advanced option in IBM StoredIQ
Data Server, is not supported.
A Connections volume that is created in IBM StoredIQ version [Link] must be fully reharvested after an
upgrade for the objects to be viewed.
Each of the subdirectories has elements under the user name directory. So, if User A created a forum, the
directory to find it is home/forums/userA/<Forum Name>. If a user created a forum inside a
community that is owned by User B, the directory to find it is home/communities/userB/<Community
Name>/forums/<Forum Name>.
For more information about Connections attributes and their use examples, see the topic about
Connections attributes in the IBM StoredIQ Data Workbench documentation.

Setting up the administrator access on Connections


IBM Connections needs an actual user account, not wasadmin, to be set up with admin and search-
admin privileges. The following procedure describes how to set up the administrator access on
Connections.
This procedure needs to be done in the WebSphere® Application Server Administrative Console by the
administrator.
1. In the Administrative Console, follow these steps.
a) Go to Users and Groups > Administrative user roles.
b) Select Add... > Administrator role.
c) Search for the Connections user account that is used to add the Connections volume in IBM
StoredIQ and add it to the role.
d) Click OK and select Save directly to the master configuration.
2. Follow these steps for each of these applications: Activities, Blogs, Communities, Dogear, Files,
Forums, News, Profiles, RichTextEditors, Search, URLPreview, and Wikis.
a) In the Administrative Console, go to Applications > Application Types > WebSphere enterprise
applications.
b) Select an application from the list.
c) Select Security role to user/group mapping > Search-admin > Map Users....
d) Search for the Connections user account that is used to add the Connections volume in IBM
StoredIQ and add it to the role.
e) Click OK > OK.
f) Select Save directly to the master configuration.
3. Follow these steps for each of these applications: Activities, Blogs, Common, Communities, Files,
Forums, Homepage, Metrics, News, Profiles, PushNotification, RichTextEditors, Search, URLPreview,
WidgetContainer, and Wikis.
a) In the Administrative Console, go to Applications > Application Types > WebSphere enterprise
applications.
b) Select an application from the list.
c) Select Security role to user/group mapping > admin > Map Users....
d) Search for the Connections user account that is used to add the Connections volume in IBM
StoredIQ and add it to the role.
e) Click OK > OK.
f) Select Save directly to the master configuration.

68 IBM StoredIQ: Data Server Administration Guide


IBM Content Manager attributes
In the SiqDocument item type, various attributes are increased when you run copy to IBM Content
Manager.
In the SiqDocument item type, the length of the following attributes is increased 128 - 256 bytes when
you run copy to IBM Content Manager:
• SiqServer
• SiqShare
• SiqInitialDirectory
• SiqFileName
• SiqContainerPath
• SiqOwner
This change is handled automatically if you do not already have an SiqDocument item type in your IBM
Content Manager server. However, if this item type exists, it must be recreated with the new attribute
lengths for this change to take effect.
Note: If you run a working CopyTo IBM Content Manager without issues or if you know that your attribute
lengths are not greater than 128 in length, then you can defer this action as you did not encounter the
attribute length issue.

Note: The attribute length is in bytes. The number of actual characters this length holds varies based on
the database code page that is used. For example, if ASCII is used, then the number of characters is equal
to the number of bytes. If UTF-8 is used, the number of bytes per character varies depending on the
characters. Without this change, you see errors if the source attributes for a copy to IBM Content Manager
exceed 128 bytes. If you see these errors, you need to take the following actions.
To extend the length of these attributes, take the following actions:
• If the SiqDocument item type does not exist in the IBM Content Manager server, create a new IBM
Content Manager volume with a CopyTo option to select SiqDocument. It creates the item type and its
attributes with the correct lengths.
• If SiqDocument item type exists in the IBM Content Manager server and you need to fix the attribute
length problem, then the administrator must delete or drop the SiqDocument item type and recreate
or update the IBM Content Manager volume that is used for copy. It automatically creates the item
types and attributes desired.
Note: Take a backup of the database before you drop and recreate SiqDocument item type. When you
drop the item type, you permanently lose all the items (documents) stored in it. If the source
documents are still available, you can run copy again to copy the data back into this item type. If no
items exist in the item type, then it is not an issue.
To drop the SiqDocument item type,
1. Delete all the items in the SiqDocument. This delete is permanent and you lose all of the data.
2. Delete the SiqDocument item type.
3. Delete all the attributes that belong to the SiqDocument
If the item type exists and contains data that you need to keep, and you need to extend these attributes,
this process is possible through direct database manipulation. However, this process is not supported and
issues that derive from it cannot be covered by IBM support. If you want this process, services must be
employed to make these database changes.

Volumes and data sources 69


Creating retention volumes
Retention volumes store data objects that are placed under retention, which means such objects cannot
be deleted for a specified period.
1. Configure your retention servers.
2. Create management or retention classes.
3. Create retention volumes.

Adding a retention volume


Retention volumes can be added and configured. Applicable volume types include CIFS (Windows
platforms) and NFS v3.
1. Go to Administration > Data sources > Volumes, and then click Retention.
2. Depending on the type of retention server you are adding, complete the fields as described in the
appropriate table.
3. Click OK to save the volume.
Note: Case-sensitivity rules apply. Red asterisks within the user interface denote required fields.

Table 24. CIFS (Windows platforms) retention volumes


Field Value Notes
Server type Select the server type.
Server Assign the server a name.
Connect as Enter the login ID.
Password Enter the password for the login
ID.
Volume Enter the name or names of the
volume to be mounted.
Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

70 IBM StoredIQ: Data Server Administration Guide


Table 24. CIFS (Windows platforms) retention volumes (continued)
Field Value Notes
Constraints Select either or both of these
options:
• Only use __ connection
process(es): Specify a limit for
the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.

Table 25. NFS v3 retention volumes


Field Value Notes
Server type Select the server type.
Server Assign the server a name.
Volume Enter the name or names of the
volume to be mounted.

Index options Select either or both of the These options are selected by
Index options checkboxes. default.
• Include system metadata for
data objects within
containers.
• Include content tagging and
full-text index.

Volumes and data sources 71


Table 25. NFS v3 retention volumes (continued)
Field Value Notes
Constraints Select either or both of these
options:
• Only use __ connection
process(es): Specify a limit for
the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and
full-text searches, you might
want to regulate the load on
the server by limiting the
harvester processes. The
maximum number of harvest
processes is automatically
shown. This maximum
number is set on the system
configuration tab.
• Control the number of
parallel data object reads:
Designate the number of
parallel data object reads.

Creating system volumes


System volumes support volume export and import. When you export a volume, data is stored on the
system volume. When you import a volume, data is imported from the system volume.
1. Go to Administration > Data sources > Specify volumes > Volumes.
2. Select the System tab, and then click Add system volumes.
3. Enter the information described in the table below, and then click OK to save the volume.
Note: Case-sensitivity rules apply. Red asterisks within the user interface denote required fields.

Table 26. System volume fields, descriptions, and applicable volume types
Field Value Applicable volume type
Server type Select the type of server. • CIFS (Windows platforms)
• NFS v2, v3

Server Enter the name of the server • CIFS (Windows platforms)


where the volume is available
• NFS v2, v3
for mounting.
Connect as Enter the logon ID used to • CIFS (Windows platforms)
connect and mount the defined
volume.
Password Enter the password used to • CIFS (Windows platforms)
connect and mount the defined
volume.

72 IBM StoredIQ: Data Server Administration Guide


Table 26. System volume fields, descriptions, and applicable volume types (continued)
Field Value Applicable volume type
Volume Enter the name of the volume to • CIFS (Windows platforms)
be mounted.
• NFS v2, v3

Constraints Only use __ connection • CIFS (Windows platforms)


process (es): Specify a limit for
• NFS v2, v3
the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and full-
text searches, you may want to
regulate the load on the server
by limiting the harvester
processes. The maximum
number of harvest processes is
automatically shown. This
maximum number is set on the
system Configuration tab.

Export and import of volume data


Metadata and full-text indexed data can be collected or exported from separate locations, such as data
servers in various offices of the enterprise. When the data is available, it can be imported to a single
location, such as a headquarter's office data server, where selected files might be retained.
Only primary and retention volume data can be exported or imported with the export and import feature.
Discovery export and system volumes cannot be imported or exported. The target location of an export or
the source location of an import is always the IBM StoredIQ system volume.
Export and import volume processes can be run as jobs in the background. These jobs are placed into
their prospective queues, and they are run sequentially. When one job completes, the next one
automatically starts. These jobs can be canceled at any time while they are running. Canceling one import
or export job also cancels all the jobs that come after the one canceled. Because the export jobs and
import jobs are in separate queues, canceling one type of job does not cancel jobs in the other queue. The
jobs cannot be restarted.

Exporting volume data to a system volume


When volume data is exported to a system volume, the export process creates two files: a binary file and
a metadata file, which contains the exported data.
The export process creates two files: a binary file and a metadata file, which contains the exported data.
These files' names contain the following information:
• Data server name and IP address
• Volume and server names
• Time stamp
Note the following information:
• The exported data consists of data from the selected volume and any related information that describes
that data except for volume-specific audits.
• The exported data must be made available to the import data server before it can be imported. It might
require you to physically move the exported data to the system volume of the import data server.
• Licenses on the import appliances are enabled automatically if a feature of the imported volume
requires it (such as Exchange licenses).

Volumes and data sources 73


1. Go to DSAdmin > Administration > Data sources > Volumes.
2. Select a volume of data to export from the list of volumes by clicking the Discovery export link in the
far right column.
3. Complete the Export volumes details window, described in this table.
4. Click OK. A dialog box appears to show that the data is being exported.
5. To monitor the export progress, click the Dashboard link. To cancel the export process, under the
Jobs in progress section of the dashboard, click Stop this job. Alternately, click OK to return to the
Volumes page.
Note: The job cannot be restarted.

Option Description
Server The name of the server where the data is.
Volume The name of the volume where the data is.
Export path (on This path is where to save the data on the system volume. The default path is /
system volume) exports. You can edit the export path. The specified location is automatically
created if necessary.
Description Enter a description of the exported data.
(optional)
Export full-text On a data server of the type DataServer - Classic, you can select this option to
index export the volume's full-text index. This option is available only if the volume
has a full-text index.
On a data server of the type DataServer - Distributed, this option is always
shown and is checked by default because the full-text index is a part of the
Elasticsearch index in general. You cannot change the setting.

Importing volume data from a system volume


An imported volume looks, acts, and is, just like a volume that is originally defined and harvested on the
data server with a few key differences.
For imported volumes, any action or relationship that is valid for a non-imported volume is valid for an
imported volume, with a few exceptions:
• Logs and audit trails that capture the activity on the volume before the import is not available. However,
the import itself is audited.
• The imported volume can be reharvested if the appliance has the proper network access and rights to
the original source server and volume.
• The imported volume can be reharvested if the data server has the proper network access and rights to
the original source server and volume.
• The data viewer works only if the appliance has the proper network access and rights to the source
server and volume. You must have access and permission on export servers and volumes if the file you
want to view was migrated to a secondary server at the time of the export.
Notes:
• When a volume with a licensed feature is imported into a data server that does not use licensing, the
license is imported along with the volume. To see the licensed features, users need to log out and then
log back in to the data server.
• Imports can only occur between the same type of data server. A volume export from a data server of the
type DataServer - Classic cannot be imported into a data server of the type DataServer - Distributed,
and vice versa.
1. Make sure that exported data file is present in the system volume of the import appliance.
2. Go to Administration > Data sources > Volumes, and then click either the Primary or Retention tab.

74 IBM StoredIQ: Data Server Administration Guide


3. Click the Import volume link at the top of either the primary or the retention volume lists. The Import
volumes page appears, listing all of the volumes available for import. By default, the data server
searches for available volumes to import in the /imports directory of the system volume. If you
placed the exported data to another path, click Change path and enter the appropriate path.
4. Click OK. The Import volumes page now displays the following information about the imported
volumes. See the following table.
5. From the list of volumes, select a volume to import by clicking the Import link in the last column on
the right.
6. Select the Import full-text index checkbox to import the selected volume's full-text index.
On a data server of the type DataServer - Classic, this option is active only if the volume has a full-text
index.
On a data server of the type DataServer - Distributed, this option is always shown and is checked by
default because the full-text index is a part of the Elasticsearch index in general. You cannot change
the setting.
7. Select the Overwrite existing volume checkbox to replace the existing data of the volume with the
imported data.
8. Click OK. A dialog box appears to show that the volume is being imported. If the volume exists and
unless the Overwrite existing volume option is selected. OK is not enabled .
9. To view import progress, click the Dashboard link in the dialog box. To cancel the import process,
under the Jobs in progress section of the dashboard, click Stop this job, or click OK to return to the
Manage volumes page.
Option Description
Server and volume The server name and volume name where the data physically is.
Description A description added when the volume was exported.
Volume type The types of volume, which are Exchange, SharePoint, and other types.
Category The category of volume, which is Primary or Retention.
Exported from The server name and IP address of the server from which the data was
exported.
Export date The day and time the data was exported.
Total data objects The total number of data objects that are exported for the volume.
Contains full-text Indicates whether the full-text index option was chosen when the data was
index exported. For the index type distributed, this field will always show the
value Yes.
Index type Indicates which type of data server generated the export. You can import
only volumes with an index type that matches the type of data server you
are working with.

Creating discovery export volumes


Discovery export volumes contain the data produced from a policy, which is kept so that it can be
exported as a load file and uploaded into a legal review tool. Administrators can also configure discovery
export volumes for managing harvest results from cycles of a discovery export policy.
1. Go to Administration > Data sources > Specify volumes, and then click Volumes.
2. Click Discovery export, and then click Add discovery export volumes.
3. Enter the information described in the table below, and then click OK to save the volume.
Note: Case-sensitivity rules for each server type apply. Red asterisks within the user interface denote
required fields.

Volumes and data sources 75


Table 27. Discovery export volumes: fields, required actions, and applicable volume types.
Field Value Applicable volume type
Server type Select the type of server. • CIFS (Windows platforms)
• NFS v2, v3

Server Enter the name of the server • CIFS (Windows platforms)


where the volume is available
• NFS v2, v3
for mounting.
Connect as Enter the logon ID used to • CIFS (Windows platforms)
connect and mount the defined
volume.
Password Enter the password used to • CIFS (Windows platforms)
connect and mount the defined
volume.
Volume Enter the name of the volume to • CIFS (Windows platforms)
be mounted.
• NFS v2, v3

Constraints Only use __ connection • CIFS (Windows platforms)


process (es): Specify a limit for
• NFS v2, v3
the number of harvest
connections to this volume. If
the server is also being
accessed for attribute and full-
text searches, you may want to
regulate the load on the server
by limiting the harvester
processes. The maximum
number of harvest processes is
automatically shown. This
maximum number is set on the
system Configuration tab.

Deleting volumes
Administrators can delete volumes from the list of available data sources, provided that the data server is
connected to the gateway.
Note the following regarding deleted volumes:
• Deleted volumes are removed from target sets.
• Deleted volumes are removed from all volume lists, both from IBM StoredIQ and IBM StoredIQ
Administrator.
• Within created jobs, steps that reference deleted volumes are implicitly removed, meaning that a job
might contain no steps. The job itself is not deleted.
• Applicable object counts and sizes within IBM StoredIQ Administrator adjust automatically.
• Object counts and sizes within user infosets remain the same. Remember, those user infosets were
created at a specific point in time when this data source was still available.
• Users who explore a specific data source and any generated reports no longer reference the deleted
volume.

76 IBM StoredIQ: Data Server Administration Guide


• No exceptions are raised on previously run actions. Instead, the data is no longer available. For
example, if an infoset is copied that contained data objects from a volume that was deleted, no
exception is raised.
• If you mark a desktop volume for deletion, it is automatically removed from the Primary volume list;
however, the status of that workstation is set to uninstall in the background. When the desktop client
checks in, it sees that change in status and uninstalls itself.
Note: If retention volumes contain data, they cannot be deleted as IBM StoredIQ is the source of record.
Instead, you see the Under Management link.
1. Go to Administration > Data sources > Specify volumes > Volumes.
2. Click the tab of the volume type that you want to delete: Primary, Retention, System, or Discovery
export.
3. Click Delete, and in the confirmation dialog window, click OK.
The volume is deleted, removing it from the list of available volumes.

Action limitations for volume types


Volume types have different action limitations.
IBM StoredIQ imposes some action limitations on volume types, which are identified in this table.

Table 28. Action limitations for volume types


Action type Source limitations Other restrictions
Copy from • Box Copy from SharePoint Online
supports copy to CIFS, FileNet,
• Chatter and SharePoint
• CIFS 2007/2010/2013/2016/Online
• CMIS target sets.
• Connections Copy from Box supports copying
files to CIFS, NFS, FileNet, and
• Desktop
Box target sets.
• Documentum
Copy from Connections supports
• Domino copying files to CIFS targets.
• Exchange
Copy from OneDrive for Business
• FileNet supports copying files to CIFS
• HDFS and NFS shares.
• IBM Content Manager
• Jive
• NewsGator
• NFS
• OneDrive
• OpenText Livelink
• SharePoint

Volumes and data sources 77


Table 28. Action limitations for volume types (continued)
Action type Source limitations Other restrictions
Copy to Primary Volume For copy policies that target
SharePoint volumes, the
• Box destination can write documents
• CIFS into an existing custom library.
Two known SharePoint library
• CMIS
configuration requirements are
• Documentum as follows.
• FileNet • For the Require documents to
• HDFS be checked out before they
• IBM Content Manager can be edited option, select
No. This option is found within
• NFS Versioning settings.
• SharePoint • Each column that is marked as
Required must specify a
default value. In the Author:
Edit Default Value dialog box,
select the Use this default
value option, and then supply a
default value. This option is
found within Common default
value settings.
Both configuration options are
available within the Library
settings of SharePoint.
For copy policies that target
FileNet volumes, creating
directories or re-creating the
directory structure is not
supported.

Copy to (Box volume) • CIFS Box volumes can be added only


from IBM StoredIQ
• NFS Administrator, not from IBM
• SharePoint StoredIQ Data Server. Preserve
owners is only supported for
SharePoint and CIFS source
volumes. Map permissions is
only supported for CIFS source
volumes.
Copy to (OneDrive volume) • CIFS OneDrive volume can be added
only from IBM StoredIQ
• NFS Administrator, not from IBM
StoredIQ Data Server. For copy
policies that target OneDrive
volumes, the destination folder
needs to be set to the name of an
existing drive at least. It cannot
be left blank.
Copy to (retention) • CIFS
• NFS

78 IBM StoredIQ: Data Server Administration Guide


Table 28. Action limitations for volume types (continued)
Action type Source limitations Other restrictions
Copy to SharePoint Online • CIFS The Copy to SharePoint Online
function does not include the
• FileNet ability to update some file system
• SharePoint metadata or custom metadata.
2007/2010/2013/2016/Online For example, the target file's
Create Time and Modified Time
reflect the time of the copy, not
the source file's Create Time and
Modified Time.
Delete • Box The Delete action removes data
objects (not directories) from the
• CIFS source volume and is limited to
• Desktop system level objects. If a
• Documentum responsive object targeted for
deletion is within a container
• HDFS object (such as a .zip or .tar file),
• NFS the container level object is the
• SharePoint object that is deleted. Thus, any
other objects within the container
file are also deleted.
A Delete action on a responsive
object attached to a CIFS
based .msg, .eml, or .pst file also
results in the container object
being deleted including any other
objects attached to or contained
within the container object.
The respective logging
information reflects the container
level object, indicating success or
failure or the action.
For desktop volumes, certain
files must not be deleted. For
more information, see “Special
considerations for desktop delete
actions” on page 98.

Volumes and data sources 79


Table 28. Action limitations for volume types (continued)
Action type Source limitations Other restrictions
Discovery export from • Box
• CIFS
• CMIS
• Connections
• Desktop
• Documentum
• Exchange
• FileNet
• HDFS
• IBM Content Manager
• Jive
• NewsGator
• NFS
• OneDrive for Business
• OpenText Livelink
• SharePoint

Discovery export to • CIFS Considered discovery export


volumes (category) not harvested
• NFS

Move from • Box Move from Box supports copying


files to CIFS, NFS, FileNet, and
• CIFS Box target sets.
• Desktop
• Documentum
• NFS

Modify security • CIFS


• NFS

Move to • CIFS For move policies that target


• Documentum FileNet volumes, creating
directories or re-creating the
• FileNet directory structure is not
• NFS supported.
• SharePoint

Volume limitations for migrations


The IBM StoredIQ supported source-infoset members for migrations to SharePoint targets have
limitations.
For migrations that preserve version hierarchies, IBM StoredIQ preserves them at the file system level. If
container-member objects are present in the source infoset, they are copied to new version chains,
independent of the source-container version. Therefore, member-version numbers might not match the
parent-container version.

80 IBM StoredIQ: Data Server Administration Guide


Data harvesting
Harvesting or indexing is the process or task by which IBM StoredIQ examines and classifies data in your
network.
Running a Harvest every volume job indexes all data objects on all volumes.
• A full harvest can be run on every volume or on individual volumes.
• An incremental harvest only harvests the changes on the requested volumes
These options are selected when you create a job for the harvest. A harvest must be run before you can
start searching for data objects or textual content. An Administrator initiates a harvest by including a
harvest step in a job.
Most harvesting parameters are selected from the Configuration subtab. You can specify the number of
processes to use during a harvest, whether a harvest must continue where it left off if it was interrupted
and many other parameters. Several standard harvesting-related jobs are provided in the system.

Harvesting with and without post-processing


You can separate harvesting activities into two steps: the initial harvest and harvest post-processing. The
separation of tasks gives Administrators the flexibility to schedule the harvest or the post-process loading
to run at times that do not affect system performance for system users. These users might, for example,
be running queries. Examples of post-harvest activities are as follows:
• Loading all metadata for a volume.
• Computing all tags that are registered to a particular volume.
• Generating all reports for that volume.
• If configured, updating tags, and creating explorers in the harvest job.

Incremental harvests
Harvesting volumes takes time and taxes your organization's resources. You can maintain the accuracy of
the metadata repository quickly and easily with incremental harvests. With both of these features, you
can ensure that the vocabulary for all volumes is consistent and up to date. When you harvest a volume,
you can speed up subsequent harvests by only harvesting for data objects that were changed or are new.
An incremental harvest indexes new, modified, and removed data objects on your volumes or file servers.
Because the harvests are incremental, it takes less time to update the metadata repository with the
additional advantage of putting a lighter load on your systems than the original harvests.
Note: Harvesting NewsGator Volumes: Since NewsGator objects are just events in a stream, an
incremental harvest of a NewsGator volume fetches only new events that were added since the last
harvest. To cover gaps due to exceptions or to pick up deleted events, a full harvest might be required.

Reharvesting
The behavior is the same for both types of data server:
On a reharvest, the metadata for a document is updated because only the latest version of the document
is considered. Therefore, the document might then no longer match previously applied filter criteria
although is it still part of the infoset.
On a reharvest, also the full-text index is updated. Any previously applied cartridges are automatically
reapplied to the latest document version to ensure that the results of any Step-up Analytics action are
still available in the full-text index. Step-up Analytics or Step-up Full-Text actions run after a reharvest
analyze and annotate the latest document version on the data source.

© Copyright IBM Corp. 2001, 2020 81


Harvest of properties and libraries
When harvesting private information, SharePoint volumes must use administrative roles for mounting
permission.
Note: Administrative permissions are required to harvest personal information, libraries, and objects that
are not designated as being visible to Everyone for user profiles.
The SharePoint volume needs to be mounted with administrative permissions. If the harvest is conducted
without administrative permissions, then any of the user profile’s properties that were marked as visible
to the category other than Everyone is not visible in results. To harvest users’ personal documents and
information, volumes that are mounted without administrative permissions must use credentials that
have full control on all SharePoint site collections. These collections are hosted by the user profile service
application.
To override this restriction, see [Link]

Lightweight harvest parameter settings


To conduct a lightweight harvest, certain configuration changes can be made.
With IBM StoredIQ, you can conduct many types of harvests, depending on your data needs. While in-
depth harvests are common, instances exist, where you need an overview of the data and a systemwide
picture of files' types and sizes. For example, at the beginning of a deployment, you might want to obtain
a high-level view of a substantial amount of data. It helps make better decisions about how you want to
handle harvesting or other policies in the future. The following section provides possible system
configurations for the system to process the volumes’ data in the quickest manner possible.

Determining volume configuration settings


To conduct a lightweight harvest, you can make various configuration changes.
Volume Details: When you configure data sources for a lightweight harvest, you do not need to include
content tagging and full-text indexes. By clearing this option, the system indexes the files’ metadata, not
the entire content of those files. The system can then run and complete harvests quickly. You can obtain
much information about file types, the number of files, the age of the files, file ownership, and other
information.
1. Go to Administration > Data sources > Specify volumes > Volumes.
2. On the Primary volume list page, click Add primary volumes or Edit to edit an existing volume.
3. Verify that all of the Index options check boxes are cleared (some are selected by default).
4. Optional: Edit the advanced details.
In some cases, you might want to reduce the weight of a full-text harvest. In these instances, you can
adjust the processing that is involved with the various harvest configuration controls.
Within volume configuration, the advanced settings are used to control what is harvested within the
volume. By harvesting only the directory structures that you are interested in, you can exercise some
control over the harvest’s weight.
Click Show Advanced Details.
• Include Directory: If you want to harvest a subtree of the volume rather than the whole volume,
then you can enter the directory here. It eliminates the harvest of objects that are not relevant to
your project.
• Start Directory and End Directory: You can select a beginning and end range of directories that are
harvested. Enter the start and end directories.
• Constraints: You can limit the files that are harvested through connection processes, parallel data
objects, or scoping harvests by extension. For example, with the Scope harvest on these volumes

82 IBM StoredIQ: Data Server Administration Guide


by extension setting, you can limit the files that you harvest by using a set of extensions. If you want
to harvest only Microsoft Office files, you can constrain the harvest to .DOC, .XLS, and .PPT files.
5. Click OK and then restart services.

Determining harvester configuration settings


When you conduct a lightweight harvest, you can make certain harvester configuration changes.
1. Determine Skip Content Processing settings.
Note: This setting is relevant only for full-text harvests.
You might have many files that are types for which you do not need the contents such as .EXE files. In
these instances, you can add these file types to the list of files for which the content is not processed.
There are two points to consider when to skip content processing:
• You do not spend time harvesting unnecessary objects, which can be beneficial from a time-saving
perspective.
• Later, you have the option of viewing the content of the skipped files. It creates more work,
reharvesting these skipped files.
2. Determine which Locations to ignore.
There might be instances where large quantities of data are contained in subdirectories, and that data
is not relevant to your harvest strategy. For example, you might have a directory with a tree of source
code or software archive that is not used as a companywide resource. In these cases, you can
eliminate these directories from the harvests by adding the directory to the Locations to ignore. These
locations are not specific to a volume, but can instead be used for common directories across
volumes.
3. Determine Limits.
• Maximum data object size: This setting is only relevant for full-text harvests. In cases many large
files, you might want to eliminate processing those files by setting the Maximum data object size
to a smaller number. The default value is 1,000,000,000. You can still collect the metadata on the
large files, so you can search for them and determine which files were missed due to the setting of
this parameter.
4. Determine Binary Processing.
If the standard processing cannot index the contents of a file, binary processing is extra processing
that can be conducted. For lightweight harvests, the Run binary processing when text processing
fails checkbox must be cleared as this setting is only relevant for full-text harvests.

Determining full-text settings


When you conduct a lightweight harvest, you can make full-text-index setting configuration changes.
1. Determine Limits.
Limit the length of words to be harvested by selecting the Limit the length of words index to __
characters option. The default value is 50, but you can reduce this number to reduce the quantity of
indexed words.
For more information about the configuration settings, see configuring Limits.
2. Determine Numbers.
If there are large quantities of spreadsheet files, you can control what numbers are indexed by the
system.
For more information about the configuration settings, see configuring Numbers.

Determining hash settings


When you conduct a lightweight harvest, you can make certain hash-setting configuration changes.
Before you change any settings, review the information in “Configuring hash settings” on page 23.

Data harvesting 83
• Determine the hash settings.
A file hash is a unique, calculated number that is based on the content of the file. By selecting Partial
data object content, you reduce the processing to create the hash. However, the two different data
objects might create the same hash. It is a small but potential risk and is only relevant for full-text
harvests.

Pausing and resuming harvests


Configure a resumable harvest to avoid complete reharvesting when a harvest is stopped and restarted.
You can pause a harvest and have it resume from the point where it was interrupted by means of a
resumable harvest. If a long running harvest that is not configured as resumable is stopped, it starts
crawling the data source all over again when the harvest is started anew. However, a resumable harvest
marks the location where it was last stopped and resumes harvesting at that location when restarted.
The feature is available only from the IBM StoredIQ Data Server Admin UI and is supported for full
harvests of CIFS and NFS data sources. It is not available on the AppStack and is not supported for
harvests of OneDrive or IBM Connections data sources.
If the data source changes after a harvest was paused, those changes might not be reflected in the
resumed harvest. To catch any such changes, run an incremental harvest after the resumable harvest is
complete. For example, if the harvest discovered the directory a and, after some time, the harvest was
paused. Then, the directory was deleted or otherwise modified. When the harvest is resumed, those
changes do not show up in the harvest because the harvest does not revisit previously harvested
locations.
1. In the IBM StoredIQ Data Server Admin UI, create a job with a Run harvest step.
2. When you configure this step, select the Resumable harvest option.
3. Start this job.
The job can be stopped at any time, which basically means the harvest is paused. Wait until the job is
completely stopped before you restart the harvest.
4. At any later time, start the job again to complete the harvest.
A minimum wait time of 2 minutes must be observed before starting the job again.
5. Run an incremental harvest to pick up any changes made after the harvest was paused.

Monitoring harvests
You might want to check large, long-running harvests regularly to get a status and to detect issues.
Full harvests of volume are the most time-consuming operations in IBM StoredIQ. Depending on the size
of the data sets, they can take days or even weeks. To get a quick status of active harvests, you can use
the Harvest Tracker tool. The tool's output can also be used for troubleshooting issues with long-running
harvests.
The tool is designed to be least intrusive with regard to data server operation.
For each check, information about all active harvests is written to a log file. At the end of a harvest, a
summary is displayed that includes findings about bottlenecks if any. By default, all the information is
written to the /deepfs/config/harvest_tracker.log file on the data server.
The records in the log file can give answers to these questions:
• Are there any extended or expired objects?
Check all entries that are labeled ObjTracker:.
• Are objects accumulating in the queues?
Compare the entries that are labeled Ingest Q: Cur len:, Output Q: Cur len:, Findex Q:
Cur len:, and TktMstr: Cur len:.

84 IBM StoredIQ: Data Server Administration Guide


• Is the object count increasing at a good pace between the checks?
A stuck object count indicates a problem with the currently processed element.
• Which files are being harvested currently?
The paths labeled File: in Ingest Q, Output Q, Findex Q, and TktMstr sections can give you a
rough idea about the part of the volume being harvested at the moment. In particular, file paths listed in
the TktMstr section indicate the most up-to-date file paths being processed.
Other observations that you might make:
• The Ingest Q queue is fuller than the others most of the time. This hints at the text extraction not
happening fast enough.
• The Output Q and Findex Q queues are almost empty. This means node-indexing and full-text
indexing aren't bottlenecks.
• Text extraction from image files is taking a long time.
• Files such as large images and complex scanned image PDF files need extended processing time and
thus slowing down the harvest.
1. Using an SSH tool, log in to the data server VM as root.
2. Run the Harvest Tracker tool.
The following command, for example, has the tool exit after 100 loops of 30 seconds, where up to 5
file names per queue are written to the output at the default location:

python32 /usr/local/storediq/bin/util/harvest_tracker.pyc –l 100 -q 5

For the detailed command syntax and more information about the records written to the log file, see
“The Harvest Tracker tool” on page 85.
You can also run the command with the -h option to display the supported options and the default
values.
At any time, you can stop the tool by pressing Enter and then entering y as confirmation.

The Harvest Tracker tool


Find details about the command syntax and the tool's output.

Command syntax
The Harvest Tracker tool is a Python script that you run on the data server. The command syntax is as
follows:
python32 /usr/local/storediq/bin/util/harvest_tracker.pyc
-lloops ,--loop= loops

-t loop_time ,--loop-time= loop_time -o logfile ,--logfile= logfile

-qqcount ,--qcount= qcount

loops
Specifies how many times the tool is to check the data server services. The default value is maxint,
which corresponds to 214,748,647.
loop_time
Specifies the time interval for the checks in seconds. The default value is 30.
logfile
Defines the path to the log file. The default log file is /deepfs/config/harvest_tracker.log

Data harvesting 85
qcount
Specifies the maximum number of objects in queue to be listed by name.
Running the command python32 /usr/local/storediq/bin/util/harvest_tracker.pyc -h displays the
supported options and the default values.
At any time, you can stop the tool by pressing Enter followed and then entering y as confirmation.
While the tool is running, you will see additional messages like this one on the terminal where you started
the tool:
Tue Feb 26 16:41:14 2019:pubsub/
[Link]:ReconnectingPBClientFactory._onRemoteOk
These messages come from the data server and cannot be suppressed. They have nothing to do with the
Harvest Tracker tool and can, therefore, be ignored. They are not written to the harvest_tracker.log
file.

Output for active harvests


The statistics are written to an individual output block for each harvest, where the blocks are separated
by dashed lines. The first line of each block shows the volume ID and the start time of that specific
harvest. The statistics include the following information:
Ingest Q: Cur len:
Number of objects in the queue that are waiting for text extraction
Ingest Q: Total in:
Total number of objected that entered the queue since the harvest started
Ingest Q: File:
Full path to the file in the queue
Output Q: Cur len:
Number of objects in the queue that are waiting for node-indexing (PostGres)
Output Q: Total in:
Total number of objects that entered the queue since the harvest started
Output Q: File:
Full path to the file present in the queue
Findex Q: Cur len:
Number of objects that are queued for full-text indexing (into Lucene)
Findex Q: Total in:
Total number of objects that entered the full-text indexing queue since the harvest started
Findex Q: File:
Full path to the file in the full-text indexing queue
TktMstr: Cur len:
Number of objects being actively tracked by Ticket Master
TktMstr: Total in:
Total number of objects tracked by Ticket Master so far
Harvest: Time to complete:
Expected remaining processing time
Harvest: Percent complete:
Percent complete
Harvest: Estimated vol size:
Estimated size of the volume being harvested
Harvest: Obj/s:
Object processing rate
Harvest: Max time file:
Name of the file taking longest processing time

86 IBM StoredIQ: Data Server Administration Guide


Harvest: Max time val:
Processing time of the file listed under Harvest: Max time file:
Harvest: Max size file:
Name of the largest file encountered so far
Harvest: Max size val:
Size of the file listed under Harvest: Max size file:
ObjTracker: VolId:
The ID of the volume being harvested
ObjTracker: Files extended:
A comma separated list of file paths to files that are taking longer than the normal processing time is
ObjTracker: Files expired:
A comma separated list of file paths to files whose processing could not be completed in the allotted
maximum time

[Thu Oct 31 20:46:38 2019] ------------------------------


[Thu Oct 31 20:47:08 2019] Harvest: ---- VolId: 260, Name: ThisVolume, Start: Thu Oct 31
20:41:23 2019 ----(active)
[Thu Oct 31 20:47:08 2019] Ingest Q: Cur len: 55, Total in: 813
[Thu Oct 31 20:47:08 2019] Ingest Q: File: d3/xlsfiles/Certificate or license [Link]
[Thu Oct 31 20:47:08 2019] Ingest Q: File: d3/xlsfiles/[Link]
[Thu Oct 31 20:47:08 2019] Ingest Q: File: d3/xlsfiles/[Link]
[Thu Oct 31 20:47:08 2019] Output Q: Cur len: 12, Total in: 1203
[Thu Oct 31 20:47:08 2019] Output Q: File: d3/xlsfiles/[Link]
[Thu Oct 31 20:47:08 2019] Findex Q: Cur len: 472, Total in: 1085
[Thu Oct 31 20:47:08 2019] Findex Q: File: d3/nsf/[Link]
[Thu Oct 31 20:47:08 2019] Findex Q: File: d3/nsf/[Link]
[Thu Oct 31 20:47:08 2019] Findex Q: File: d3/nsf/[Link]
[Thu Oct 31 20:47:08 2019] Tkt Mstr: Cur len: 34, Total in: 112
[Thu Oct 31 20:47:08 2019] Tkt Mstr: File: d3/pptfiles/Device identifier or [Link]
[Thu Oct 31 20:47:08 2019] Tkt Mstr: File: d3/pptfiles/Discharge [Link]
[Thu Oct 31 20:47:08 2019] Tkt Mstr: File: d3/pptfiles/Relative's full [Link]
[Thu Oct 31 20:47:08 2019] Harvest: Time to complete: 1 minute 29 seconds
[Thu Oct 31 20:47:08 2019] Harvest: Percent complete: 78.01
[Thu Oct 31 20:47:08 2019] Harvest: Estimated vol size: 773 objects
[Thu Oct 31 20:47:08 2019] Harvest: Obj/s: 3.03
[Thu Oct 31 20:47:08 2019] Harvest: Max time file: d1/pst/[Link], Max time val: 276.00
[Thu Oct 31 20:47:08 2019] Harvest: Max size file: sips/bigsip_58m.pdf, Max size val: 58616009
[Thu Oct 31 20:47:08 2019] ObjTracker: files extended: []
[Thu Oct 31 20:47:08 2019] ObjTracker: files expired: []
[Thu Oct 31 20:47:08 2019] ------------------------------

Output for finished harvests


After a harvest is complete, the tool can provide a summary of the harvest operation. The summary
contains the following information:
Ingest Q: Longest Q len:
Largest number of objects in the queue that waited for text extraction
Output Q: Longest Q len:
Largest number of objects in the queue that waited for node-indexing (PostGres)
Findex Q: Longest Q len:
Largest number of objects in the queue that waited for full-text indexing (into Lucene)
TktMstr: Longest Q len:
Maximum number of objects tracked
Harvest: Min Obj/s:
Lowest object processing rate
Harvest: Max Obj/s:
Highest object processing rate
Harvest: Max time file:
Name of file that took the longest processing time

Data harvesting 87
Harvest: Max time val:
Actual processing time of the file listed under Harvest: Max time file:
Harvest: Max size file:
Name of largest file processed
Harvest: Max size val:
Size of the file listed under Harvest: Max size file:
Harvest: Total:
Total harvest time
ObjTracker: VolId:
The ID of the volume ID that was harvested
ObjTracker: files extended:
A comma separated list of file paths to files that took longer than the normal processing time
ObjTracker: files expired:
A comma separated list of file paths to files whose processing could not be completed in the allotted
max time
Object analysis: Total obj types:
Total number of different file extensions encountered so far
Object analysis: Total obj count:
Total number of system level objects processed so far. Note that this count currently does not include
the count of objects in a container.
Object analysis: Longest tkt time:
Longest life of a ticket tracked by Ticket Master
Object analysis: ext:
Extension (type) of object
Object analysis: Count:
Number of tickets of extension/type processed
Object analysis: Max tkt time:
Longest life of ticket of this extension/type tracked by Ticket Master

[Thu Oct 31 20:48:38 2019] ------------------------------


[Thu Oct 31 20:48:38 2019] Harvest_Tracker Summary...
[Thu Oct 31 20:48:38 2019] Harvest: ---- VolId: 260, Name: ThisVolume, Start: Thu Oct 31
20:41:23 2019 ----(stale)
[Thu Oct 31 20:48:38 2019] IngestQ: Longest Q len: 813
[Thu Oct 31 20:48:38 2019] OutputQ: Longest Q len: 1324
[Thu Oct 31 20:48:38 2019] FindexQ: Longest Q len: 1196
[Thu Oct 31 20:48:38 2019] TktMstr: Longest Q len: 112
[Thu Oct 31 20:48:38 2019] Harvest: Min Obj/s: 0.01, Max Obj/s: 3.63
[Thu Oct 31 20:48:38 2019] Harvest: Max time file: d1/pst/[Link], Max time val: 282.00
[Thu Oct 31 20:48:38 2019] Harvest: Max size file: sips/bigsip_58m.pdf, Max size val: 58616009
[Thu Oct 31 20:48:38 2019] Harvest: total: 0d 0h:7m:15s
[Thu Oct 31 20:48:38 2019] ObjTracker: files extended: []
[Thu Oct 31 20:48:38 2019] ObjTracker: files expired: []
[Thu Oct 31 20:48:38 2019] Object analysis:
[Thu Oct 31 20:48:38 2019] Total obj types: 22, total obj count: 793, Longest tkt time:
32.19:
[Thu Oct 31 20:48:38 2019] ext: 'xls': count: 193 (24.00%), Max tkt time: 2.39
[Thu Oct 31 20:48:38 2019] ext: 'ppt': count: 191 (24.00%), Max tkt time: 4.91
[Thu Oct 31 20:48:38 2019] ext: 'doc': count: 185 (23.00%), Max tkt time: 32.19
[Thu Oct 31 20:48:38 2019] ext: 'pdf': count: 181 (22.00%), Max tkt time: 15.01
[Thu Oct 31 20:48:38 2019] ext: 'mail': count: 11 (1.00%), Max tkt time: 0.03
[Thu Oct 31 20:48:38 2019] ext: 'msg': count: 7 (0.00%), Max tkt time: 0.02
[Thu Oct 31 20:48:38 2019] ext: 'eml': count: 4 (0.00%), Max tkt time: 0.00
[Thu Oct 31 20:48:38 2019] ext: 'mbx': count: 3 (0.00%), Max tkt time: 0.00
[Thu Oct 31 20:48:38 2019] ext: 'note': count: 2 (0.00%), Max tkt time: 0.00
[Thu Oct 31 20:48:38 2019] ext: 'pst': count: 2 (0.00%), Max tkt time: 0.00
[Thu Oct 31 20:48:38 2019] ext: 'rar': count: 2 (0.00%), Max tkt time: 0.00
[Thu Oct 31 20:48:38 2019] ext: 'nsf': count: 2 (0.00%), Max tkt time: 0.00
[Thu Oct 31 20:48:38 2019] ext: 'jpg': count: 1 (0.00%), Max tkt time: 0.01
[Thu Oct 31 20:48:38 2019] ext: 'rtf': count: 1 (0.00%), Max tkt time: 0.00
[Thu Oct 31 20:48:38 2019] ext: 'h': count: 1 (0.00%), Max tkt time: 0.00

88 IBM StoredIQ: Data Server Administration Guide


Job configuration
In IBM StoredIQ Data Server, configure and run jobs with different functions. Jobs start tasks such as
harvests or maintenance and cleanup.
You can run jobs at the time of creation or schedule them to run at a designated future time and at regular
intervals. Jobs consist of either a single step or a series of steps. The actions available at each step
depend on the type of job. A set of predefined jobs that are ready for immediate use are provided with the
product. In addition to these jobs, you can create your own jobs. You can modify any job that is stored in
the Workspace folder or any subfolder to it. You can also delete jobs from these folders when you no
longer need them.

Table 29. Predefined jobs


Job Description
CIFS/NFS retention An unscheduled, one-step job that harvests the Windows Share/NFS retention
volume deleted files volumes, looking for files that require removal because the physical file was
synchronizer deleted from the retention file system. This job is in the Library/Jobs folder.
Database compactor A scheduled job that helps to limit "bloat" (unnecessary storage usage) in the
database. While this job runs, it must have exclusive, uninterrupted access to
the database. Administrators can override this job by logging in and then
proceed to use the system. This job is in the Library/Jobs folder.
Harvest every volume An unscheduled, one-step job that harvests all primary and retention volumes.
This job is in the Workspace/Templates folder.
System maintenance A multistep job in the Library/Jobs folder. The system is configured to run a
and cleanup system maintenance and clean up job once a day and includes these items:
• Email users about reports
• Email Administrators about reports
• Delete old reports
• Delete old harvests
• Load indexes
• Optimize full-text indexes

Update age explorers A one-step job in the Library/Jobs folder that recalculates these items:
• Owner Explorer data for Access Date
• Owner Explorer data for Modified Date
• Created Date (API only) values

Creating a job
Create jobs for different purposes, for example, to run tailored harvest or to discover retention volumes.
You can create custom jobs in the Workspace folder or in any subfolder to this folder.
To create a job:
1. Navigate to the Workspace folder on the Folders tab.
2. Select New > Job.
3. Enter a unique job name.
4. From the Save in list, select the appropriate folder, and click OK.

© Copyright IBM Corp. 2001, 2020 89


The job is created.
5. To view the job, edit its details and add steps, click Yes when prompted.
6. To set a schedule for the job, click Edit job details.
You can specify the time, date, and frequency for the job to run.
• Enter the time that the job must start, or click Now to populate the time field with the current time. If
you do not specify all of the job steps, you might want to add some time.
• Enter the date on which to run the job, or click Today to populate the date field with the current date.
• Specify how often the job must run. If you select None for the frequency, the job runs once, at the
time and date provided.
You can change these settings anytime.
7. To add steps, click Add step and select a step type from the list.
8. Configure options for the selected step type.
• Run harvest step: on the Specify harvest and load options page, configure the following options:
– Select the volumes to be harvested.
– Select the harvest type. You can choose to run a full harvest or an incremental harvest. With a
full harvest, all data objects on the selected volume are indexed. With an incremental harvest,
only files or data objects that changed since the last harvest are indexed. Incremental harvest is
the default setting.
– Schedule harvest and load. To limit resource use, you can separate harvest and load processes.
Select one of the options:
- Run the harvest and load indexes when the harvest completes.
- Run the harvest and delay the index loading to run with the next system-services job after the
harvest is completed. The system-services job is scheduled to run at midnight by default.
- Run only the harvest. Select this option if you plan to load harvested data into indexes later.
- Load the indexes only. Select this option to load previously harvested data into indexes.
– Configure harvest sampling if you want to limit the harvest to a smaller sample.
– Limit the harvest by time or total number of data objects. Enter the number of minutes or the
number of data objects.
• Discover Retention volumes step:
– From the Discover Retention volume list, select the retention volume to be used for this job.
– Specify for how long the harvest is to run.
– Limit the number of data objects to be harvested.
You can edit or remove steps at any time.
9. Click OK.

Starting a job
Run predefined or custom jobs as required.
Predefined jobs except for the Harvest every volume job are stored in the Library/Jobs folder on the
Folders tab. You can find the Harvest every volume job in the Workspace/Templates folder. Custom
jobs are stored in the Workspace folder or in any subfolders to the Workspace folder that an
administrator created.
1. Navigate to the folder that contains the job.
2. To start a job immediately:
• Click the name of the job, and in the Job details page, click Start job.

90 IBM StoredIQ: Data Server Administration Guide


When you click Edit job details, you can set or change the schedule for this job before you start it.
• Right-click the job and select Start.
In the Job Details area, the Schedule entry changes to This job is running now

Monitoring processing
You can track the system’s processing on your harvest/policy and discovery export tasks with the View
cache details feature. The appliance gathers data in increments and caches the data as it gathers it. If a
collection is interrupted, the appliance can resume collection at the point that it was interrupted, instead
of starting over from the beginning of the task.
1. From Administration > Dashboard, in the Appliance status pane, click View cache details.
2. To see the progress of a harvest or a policy, click the Volume cache tab. Or, to see discovery export job
progress, click the Discovery export cache tab.
Information for a job is only available while the job is running. After a task is completed, the job
disappears from the list.

Table 30. Harvest/Volume cache details: Fields, descriptions, and values


Field Description Value
Name The name of the volume that is
being harvested.
Start date The time that the job started.
Type The type of job that is run. • Copy
• Harvest - full
• Harvest - incremental

State The status of the process. • Caching: The volume cache is


being created/updated by a
harvest or policy
• Cached: Creation or update of
volume cache is complete
(harvest only)
• Loading: Volume cache
contents are being
successfully loaded into the
volume cluster

Full-text It indicates whether a full-text Yes or No


harvest is being conducted.
View audit link details Link to the harvest/policy audit
page.

Table 31. Discovery export cache details: Fields, descriptions, and values
Field Description Value
Name The name of the volume that is
being processed.
Start date The starting date/time for the
process.

Job configuration 91
Table 31. Discovery export cache details: Fields, descriptions, and values (continued)
Field Description Value
Type Type of file that is being Discovery export
prepared for export.
State The status of the discovery • Aborted: Discovery export
export job. policy was canceled or
deleted by the user.
• Caching: The volume cache is
being created/updated by a
harvest or policy.
• Cached: Creation or update of
volume cache is complete
(harvest only).
• Loading: Volume cache
contents are being
successfully loaded into the
volume cluster.

Full-text Whether a full-text harvest is Yes or No


being conducted.

Determining whether a harvest is stuck


The speed of a harvest depends on volume size and processing speed; however, harvests do occasionally
become stuck and are unable to complete successfully. Use the procedures that are outlined here to
troubleshoot the harvest process.
1. Click Administration > Dashboard > Jobs in Progress to verify that your job continues to run.
2. In Jobs in Progress, note the Total data objects encountered number.
3. Wait 15 minutes, leaving the harvest to continue to run.
4. Note the new value Total data objects encountered, and then compare it to that value denoted
previously.
5. Answer the questions in this table.
Option Description
Question Action
Question 1: Is the Total • Yes: If the number of encountered data objects continues to increase,
data object encountered then the harvest is running correctly.
counter increasing?
• No: If the number of encountered objects remains the same, then go to
Question 2.

Question 2: Is the load To view load averages, on Appliance status > About appliance > View
average up? details > System services, look at the load averages in the Basic system
information area.
• Yes: If the load averages number is up, the harvest might be stuck. Call
technical support to report that the harvest is stuck on files.
• No: The job is not really running. It means that the job must be
restarted. Go to Question 3.

Question 3: Did the job • Yes: If the job completed successfully after it was restarted, then the
complete on the second harvest is not stuck.
pass?

92 IBM StoredIQ: Data Server Administration Guide


Option Description

• No: The job did not complete successfully. Call technical support to
report a job that does not complete.

Job configuration 93
Desktop collection
IBM StoredIQ Desktop Data Collector (also referred to as desktop client) enables desktops as a volume
type or data source, allowing them to be used just as other types of data sources. IBM StoredIQ Desktop
Data Collector can collect PSTs, compressed files, and other data objects.
After the desktop client is installed on a desktop, you connect and register it with the data server. That
desktop is available as a data source within the list of primary volumes. Additionally, while the snippet
support and the Step-up Snippet action are supported by IBM StoredIQ Desktop Data Collector, a
desktop cannot be the target or destination of an action.
The data server creates and maintains an index for the desktop volume. Metadata and full-text indexes
are supported. A harvest gathers file system metadata, decrypts, cracks file type, extracts text and
content metadata from local files. After the desktop data is indexed, the data can be searched and acted
upon even while desktops are offline or unreachable.
Viewing desktop data is not possible in IBM StoredIQ Data Workbench. Content preview for desktop
documents is possible in IBM StoredIQ Insights if the documents are full text indexed and the volumes
are managed by a data server of the type DataServer - Distributed.

IBM StoredIQ Desktop Data Collector client prerequisites


The IBM StoredIQ Desktop Data Collector agent works with the following operating systems:
• Windows 7 32- and 64-bit
• Windows 8 32- and 64-bit
• Windows 10 32- and 64-bit
• Windows Server 2003, 2008, 2012, 2016
Note: For Desktop Agent with Windows Vista SP2 or Windows Server 2008 SP2, you must use Service
Pack 2 and [Link] It is a required Microsoft update.
Installation requires administrative privileges on the desktop. Before you use IBM StoredIQ Desktop Data
Collector, you need to notify users that desktop collection is going to be conducted and make them aware
of the following items:
• The desktop must be connected over the network during data collection. If the connection is
interrupted, IBM StoredIQ Desktop Data Collector resumes its work from the point at which it stopped.
• Users might notice a slight change in performance speed, but that they can continue working normally.
Desktop collection does not interfere with work processes.
• Certain actions can be taken from the tray icon: Right-click for About, Restart, Status, and Email Logs
(which packages logs in to single file and starts the email client so that the user can mail them to the
IBM StoredIQ administrator).
To ensure the network access for desktop volumes, the following port ranges must be open through a
firewall.
• 21000-21004
• 21100-21101
• 21110-21130
• 21200-21204
All communications are outbound from the client. The appliance never pushes data or requests to the
desktop. The IBM StoredIQ Desktop Data Collector pings the appliance about every 60 seconds, and the
Last Known Contact time statistic is updated approximately every 30 minutes. Additionally, IBM
StoredIQ Desktop Data Collector checks for task assignments every 5 minutes.

94 IBM StoredIQ: Data Server Administration Guide


One can download the installer application from the application in the Configuration tab. Also, the
Administrator can temporarily disable the client service on all desktops that are registered to the data
server from the Configuration tab.

Downloading the IBM StoredIQ Desktop Data Collector installer


The IBM StoredIQ Desktop Data Collector installer can be downloaded from the application.
Open ports for desktop client access to the data server on OVA deployed systems. How to do this is
described in Open ports for desktop client access to the data server.
1. Enter the https:// IP address of the data server in the browser.
2. Log in with administrator credentials.
3. Go to Administration > Configuration.
4. On the System Configuration page under Application, click Desktop Settings.
5. Click Download the desktop client installer and save the [Link] file to your local file
system.
IBM StoredIQ Desktop Data Collector can be installed only on a Windows workstation.

IBM StoredIQ Desktop Data Collector installation methods


The IBM StoredIQ Desktop Data Collector is provided as a standard MSI file and is installed according to
the typical method (such as Microsoft Systems Management Server (SMS )) used within your organization.
During installation, the host name and IP address of the IBM StoredIQ data server must be supplied. If
the installation is conducted manually by users, you must provide this information to them using email, a
text file, or another method.
IBM StoredIQ Desktop Data Collector can be installed with the following methods.

Mass distribution method (SMS)


The appliance ID is part of the distribution configuration. This method supports passing installation
arguments as MSI properties.
• Required
– SERVERACTIONNODEADDRESS IP address or host name for the Action node. When the installation
is not silent, the user is prompted for IP address or host name. The default is the value of this
argument. This field must be entered accurately or manual correction is required in the desktop
configuration file.
• Optional
– SERVERACTIONNODEPORT Port number for the Agent on the Action node. Defaults to 21000, and
can be only changed when the agent connects on a different port that is then mapped to 21000.
– NOTRAYICON Specifies whether the agent displays the IBM StoredIQ Desktop Data Collector tray
icon while it is running. Changing the setting to 1 forces the agent to run silently and not display a tray
icon.
– SERVERACTIONNODEADDRESS IP address or host name for the Action node. When the installation
is not silent, the user is prompted for IP address or host name. The default is the value of this
argument. This field must be entered accurately or manual correction is required in the desktop
config file.
• Emailing links Send a link within an email such as file:\\g:\group\install\Client-
[Link]. The link can be to any executable file format such as .BAT, .VBS, or .MSI. The .BAT/.VBS
formats can be used to pass client arguments to an .MSI file. The user who clicks the link must have
administrative privileges.
• NT Logon Script A .BAT file or .VBS script starts msiexec. Examples are given here:

Desktop collection 95
– /i: Install
– /x {7E9E08F1-571B-4888-AC08-CEA8A076F5F9}: Uninstall the agent. The product code must
be present.
– /quiet: install/uninstall runs silently. When you specify this option, SERVERACTIONNODEADDRESS
must be supplied as an argument.

Set WshShell = CreateObject("[Link]")

[Link] "%windir%\System32\[Link] /i G:\group\install\


[Link]
NOTRAYICON=0 SERVERACTIONNODEADDRESS=[Link] /q"Set WshShell
= CreateObject("[Link]")

Set WshShell = Nothing

• MSI batch file

msiexec /i G:\group\install\[Link] NOTRAYICON=1


SERVERACTIONNODEADDRESS=[Link] /q

• MSI interactive installation


The user must have administrative privileges. Any previously installed desktop client installed must be
uninstalled before a new desktop client can be installed.
The installation can be started by either double-clicking the .msi file or by right-clicking the file and
selecting Install > Run.
After the installation is complete, the desktop is available as a data source within the list of primary
volumes in IBM StoredIQ Data Server and in IBM StoredIQ Administrator.

Removing IBM StoredIQ Desktop Data Collector


If you delete a desktop volume, the respective client automatically uninstalls itself on the next check for a
pending job. Users can also manually uninstall the IBM StoredIQ Desktop Data Collector client from the
workstation.

Configuring desktop collection settings


Enable or disable desktop collection, configure upgrade settings for the desktop client, and enable
decryption of files encrypted with the Windows Encrypting File System (EFS).
1. Go to Administration > Configuration > Application > Desktop settings.
2. Enable or disable desktop services.
By default, desktop services are enabled. If you change the setting, click Apply to have the changes
take effect. Disabling the desktop services does not uninstall the desktop client from any user
workstation.
3. Configure upgrade settings for the desktop clients registered with this data server.
You can choose between automatic and manual upgrades. In the Upgrades
a) For automatic upgrades, select Upgrade all workstations from the Automatic upgrades options.
The default setting is Upgrades disabled, which means any upgrades must be done manually.
b) Select how new versions of the desktop client are made available for download. The currently
published version (the version currently available for download) is listed next to the Available
versions option. Select one of these options:
• If you want to publish a new version manually, select Manually publish new version and also
select a version.
• If you want always the latest version to be automatically available for download, select
Automatically publish the latest version.

96 IBM StoredIQ: Data Server Administration Guide


The latest client version is not necessarily the same as the currently published version. The latest
client version available is listed under Available versions. If you need the latest client, you might
need to publish, download, and install a newer client version.
c) To publish the new version (make it available for download), click Apply.
Publishing a newer client version results in the permanent removal of all older versions. Therefore,
do not publish a new version unless you are certain that the old client version is no longer needed.
4. To enable decryption of files saved in the Encrypting Files System (EFS), install a domain recovery
agent certificate or a user credential certificate.
a) In the Encrypting File System recovery agent users, click Add Encrypting File System user
b) In the dialog box, provide the following information:
• Select a Personal Information Exchange Format (PFX) file to upload. Click browse to navigate to
the file that you want to upload.
• Enter the password that protects the .PFX file.
• Enter the user name of the EFS user to whom this recovery agent belongs. The user name must
be a SAM compatible/NT4 Domain name-style user name, for example, MYCOMPANY
\esideways.
If the computer is not part of a domain and is running any version of Windows earlier than 7.0, the
user name must be the user name. If the computer is not part of a domain and is running
Windows 7.0 or later, the user name must be the name of the PC and the domain.
• Enter the password for the EFS user.
• Enter a description.
• Click OK.
The file is uploaded and the user is added to the list of recovery agent users. At any time, you can
edit or delete a user's entry.
c) Have those users for whom you installed certificates restart IBM StoredIQ Desktop Data Collector
before you start harvesting.

Configuring desktop collection


After a desktop is added as a data source, you can edit the volume configuration to adjust several
settings. Then, create and run a job with a harvest step.
In IBM StoredIQ Data Server, you can configure several options for a desktop volume that affect harvests
and actions.
When your volume definition is up-to-date, set up and run a harvest job to start collecting desktop data.
You can also trigger a harvest from IBM StoredIQ Administrator. In this case, you can only choose
whether you want to schedule the harvest or run it immediately and if you want to do a full or an
incremental harvest.
1. To change index settings and to scope the harvest, edit the volume definition.
a) Go to Administration > Data sources > Specify volumes > Volumes.
b) From the primary volume list, select the desktop volume and click Edit.
c) Adjust the index options as required.
d) To scope harvests, click Show advanced options.
Now, you can specify a regular expression to be applied to the subdirectories of the initial directory.
Folders that match the regular expression are in scope for harvests.
You can also scope harvests by file extensions. You can choose between including and excluding
data objects with the specified file extensions.
e) Click OK to save your settings.

Desktop collection 97
2. Go to Folders, click New and select Job.
3. Name the job and select the folder where the job is to be stored.
4. Select to view the job when prompted.
5. To configure a schedule for the job, edit the job details.
6. Add a step of the type Run harvest.
Configure harvest and load options for the step:
a) Select the volume that you want to harvest.
Desktop volumes appear in the list in the format hostname:workstationID, for example,
[Link]:FADBAE872A7DAF7A833040327AB8FDCDF1F6B409.
b) Set any other harvest options as required.
c) Click OK to save the settings.
7. Optional: Start the job.
You can monitor the job progress on the IBM StoredIQ Data Server dashboard.

Special considerations for desktop delete actions


Use caution when you apply any delete action to desktop data. Some files must not be deleted.
When you use the IBM StoredIQ to delete files from a desktop, they are removed permanently. They are
not transferred to the appliance or backed up to any other location. You must carefully review the infoset
of affected data objects before you take a delete action. Your organization can use custom applications or
other files that you might not want to delete. In reviewing the returned list, do not allow the following files
to be deleted.
• Anything in the c:\Windows directory
• These files in the c:\Documents and Settings\username directory:
– \UserData\*.xml
– \Cookies\*.txt
– \Start Menu\Programs\*.lnk
• Executable routines: files with the extensions .dll, .exe, and .ocx
• Drivers: files with the extensions *.sys, *.inf, and *.pnf
• Installers: files with the extensions .msi and .mst
• Important data files: files with the extensions *.dat, *.ini, *.old, and *.cat
• The following files:
– [Link]
– [Link]
– [Link]
– [Link]
– [Link]

98 IBM StoredIQ: Data Server Administration Guide


Audits and logs
The following section describes the audit and log categories in the system, including descriptions of the
various audit types and how to view and download details.

Harvest audits
Harvest audits provide a summary of the harvest, including status, date, duration, average harvest speed,
and average data object size. They can be viewed in two ways: by volume name or by the date and time of
the last harvest.
Data objects can be skipped during a harvest for various reasons such as the unavailable object or a
selected user option that excludes the data object from the harvest. The Harvest details page lists all
skipped data objects that are based on file system metadata level and content level.
All skipped harvest-audit data and other files that are not processed can be downloaded for analysis.

Table 32. Harvest audit by volume: Fields and descriptions


Harvest audit by
volume field Description
Server The server name.
Volume The volume name.
Harvest type The type of harvest that is conducted: Full Harvest, ACL only, or Incremental.
Last harvested The date and time of the last harvest.
Total system data The total number of system data objects encountered.
objects
Data objects fully The number of data objects that were fully processed.
processed
Data objects previously The number of data objects that were previously processed.
processed
Processing exceptions The number of exceptions that are produced during processing.
Binary processed The number of processed binary files.
Harvest duration The length of time of the harvest’s duration.
Status The harvest’s status: Complete or Incomplete.
Average harvest speed The average harvest speed, which is given in terms of data objects that are
processed per second.
Average data object size The average size of encountered data objects.

Table 33. Harvest audit by time: Fields and descriptions


Harvest audit by time
field Description
Harvest start The time and date at which the harvest was started.
Harvest type The type of harvest that is conducted: Full Harvest, ACL only, or
Incremental.

© Copyright IBM Corp. 2001, 2020 99


Table 33. Harvest audit by time: Fields and descriptions (continued)
Harvest audit by time
field Description
Total system data The total number of system data objects that were found.
objects
Data objects fully The total number of system data objects that were fully processed.
processed
Data objects previously The total number of system data objects that were previously processed.
processed
Processing exceptions The total number of encountered processing exceptions.
Binary processed The total number of processed binary files.
Harvest duration The length of time of the harvest’s duration.
Status The harvest’s status: Complete or Incomplete.
Average harvest speed The average harvest speed, which is given in terms of data objects that are
processed per second.
Average data object size The average size of encountered data objects.

Table 34. Harvest overview summary options: Fields and descriptions


Harvest overview
summary options field Description
Harvest type The type of harvest: Full Harvest, ACL only, or Incremental.
Harvest status The harvest's status. Options are Complete or Incomplete.
Harvest date The date and time of the harvest.
Harvest duration This duration is the length of time of the harvest's duration.
Average harvest speed The average harvest speed, which is given in terms of data objects that are
processed per second.
Average data object size The average size of encountered data objects.

Table 35. Harvest overview results options: Fields and descriptions


Harvest overview
results options field Description
Total system data The total number of system data objects that were found.
objects
Total contained data The total number of contained data objects.
objects
Total data objects The total number of encountered data objects.

Table 36. Harvest overview detailed results: Fields and descriptions


Harvest overview
detailed results field Description
Skipped - previously The number of skipped objects that were previously processed.
processed

100 IBM StoredIQ: Data Server Administration Guide


Table 36. Harvest overview detailed results: Fields and descriptions (continued)
Harvest overview
detailed results field Description
Fully processed The number of fully processed data objects.
Skipped - cannot access The number of data objects that were skipped as they might not be accessed.
data object
Skipped - user The number of data objects that were skipped because of their user
configuration configuration.
Skipped directories The number of data objects in skipped directories.
Content skipped - user The number of data objects where the content was skipped due to user
configuration configuration.
Content type known, The number of data objects for which the content type is known and partial
partial processing processing is complete.
complete
Content type known, but The number of data objects for which the content type is known, but an error
error processing content was produced while processing content.
Content type known, but The number of data objects for which the content type is known, but the
cannot extract content content might not be extracted.
Content type unknown, The number of data objects for which the content type is unknown and is not
not processed processed.
Binary text extracted, The number of data objects for which the binary text is extracted and full
full processing complete processing is completed.
Binary text extracted, The number of data objects for which the binary text is extracted and partial
partial processing processing is completed.
complete
Error processing binary The number of data objects for which an error was produced while binary
content content is processed.
Total The total number of data objects.

Viewing harvest audits


Harvest audits can be viewed from the Audit tab.
1. Go to Audit > Harvests > View all harvests. The Harvest audit by volume page opens, which lists
recent harvests and includes details about them.
2. In the Volume column, click the volume name link to see the harvest audit by time page for that
particular volume. The harvest audit by time page lists all recent harvests for the chosen volume and
includes details about each harvest.
3. In the Harvest start column, click the harvest start time link to see the harvest overview page for the
volume. You can also access the page by clicking the Last harvested time link on the Harvest audit by
volume page. The Harvest overview page provides the following options.
• Summary: Harvest type, status, date and time, duration, average harvest speed, and average data
object size.
• Results: Total system data objects, total contained data objects, total data objects.
• Detailed results: Skipped - previously processed; fully processed; skipped - cannot access data
object; skipped - user configuration; skipped directories; content skipped - user configuration;
content type known, partial processing complete; content type known, but error processing
content; content type known, but cannot extract content; content type unknown, not processed;

Audits and logs 101


binary text extracted, full processing complete; binary text extracted, partial processing complete;
error processing binary content; error-gathering ACLs; and total.
4. To view details on data objects, click the link next to the data objects under Detailed results.
• With the exceptions of skipped - previously processed, fully processed, and the total, all other
results with more than zero results have links that you can view and download results.
• The skipped data object list includes object name, path, and reason skipped. Data objects can be
skipped at the file system metadata level or at the content level. Data objects skipped at the
content level are based on attributes that are associated with the data object or its contents.
Skipped Data Objects Results Details provides details about skipped data objects.
If data objects were not harvested, you might want to download the data object's harvest audit list
details for further analysis.

Downloading harvest list details


Harvest list details can be downloaded in a .CSV format.
1. From the Harvest details page, click the active link next to the data objects under Detailed results. A
page named for the detailed result chosen (such as Skipped - user configuration or binary text
extracted, full processing complete) appears.
2. Click the Download list in CSV format link on the upper left side of the page. A dialog informs you that
the results are being prepared for download.
3. Click OK. A new dialog appears, prompting you to save the open or save the .CSV file. Information in
the downloaded CSV file includes:
• Object name
• System path
• Container path
• Message explaining why data object was skipped
• Server name
• Volume name

Import audits
Volume-import audits provide information about the volume import. This information includes the
number of data objects that are imported, the system that is exported from, the time and date of the
volume import, whether the imported volume overwrote an existing volume, and status. The volume
name links to the Import details page.

Table 37. Imports by volumes details: Fields and descriptions


Imports by volumes
details field Description
Volume The name of the imported volume.
Exported from The source server of the imported volume.
Import date The date and time on which the import occurred.
Total data objects The total number of imported data objects.
imported
Overwrite existing If the import overwrote an existing volume, the status is Yes. If the import did
not overwrite an existing volume, the status is No.
Status The status of the import: Complete or Incomplete.

102 IBM StoredIQ: Data Server Administration Guide


To view audit details of volume imports, go to Audit > Imports, and click View all imports. Then, click a
volume name in the Volume column..

Event logs
Event logs captures every action that is taken by the system and its users. It documents actions that
succeed and fail.
These actions include creating draft and published queries and tags, running policies, publishing queries,
deleting objects, configuring settings, and any other action that is taken through the interface. A detailed
list of log entries is provided in the event log messages.
You can view event logs for the current day or review saved logs from previous days, and up to 30 days
worth of logs can be viewed through the interface. If you select and clear a day of logs, those logs are
removed from the system.

Viewing event logs


Event logs can be viewed from the Dashboard or from the Audit tab.
1. Conduct either of the following actions:
a) Click Administration > Dashboard, and then locate the Event log section on the dashboard. The
current day’s log displays there by default.
b) Click the Audit tab and locate the Event logs section.
2. To view a previous day's log on the dashboard, use the View all event logs list to select the day for
which you want to view an event log.
3. Select a different day from the view event log from the list. This menu displays the event log dates for
the past 30 days. Each log is listed by date in YYYY-MM-DD format.

Subscribing to an event
You can subscribe to and be notified of daily event logs.
1. Go to Audit > Event logs.
2. Click View all event logs, and the Event log for today page opens.
3. To the right of the event log to which you want to subscribe, click Subscribe. The Edit notification page
appears.
4. In Destination, select the method by which you want to be notified of this event log. If you select
Email address, be certain to use commas to separate multiple email addresses.
5. Click OK.
Note: You can also subscribe to an event on the Dashboard. In the Event log area, click Subscribe to
the right of the event.

Clearing the current event log


An event log can be cleared from the Dashboard.
On the Administration > Dashboard, click Clear for the current view.

Downloading an event log


Event logs can be downloaded and saved.
1. When you view an event log, click the Download link for saving the data to a text file.
2. Select to save the file from the prompt. Enter a name and select a location to save the file.

Audits and logs 103


Event log messages
The following sections contain a complete listing of all ERROR, INFO, and WARN event-log messages that
appear in the Event Log of the IBM StoredIQ console.
• “ERROR event log messages” on page 104
• “INFO event log messages” on page 112
• “WARN event log messages” on page 120

ERROR event log messages


The following table contains a complete listing of all ERROR event-log messages, reasons for occurrence,
sample messages, and any required customer action.

Table 38. ERROR event log messages. This table lists all ERROR event-log messages.
Event
number Reason Sample message Required customer action
1001 Harvest was unable to open a Harvest could not allocate Log in to UTIL and restart
socket for listening to child listen port after <number> the application server.
process. attempts. Cannot kickstart Restart the data server.
interrogators. (1001) Contact customer support.
9083 Unexpected error while it is Exporting volume Contact Customer Support.
exporting a volume. 'dataserver:/mnt/demo-A'
(1357) has failed (9083)
9086 Unexpected error while it is Importing volume Contact Customer Support.
importing a volume 'dataserver:/mnt/demo-A'
(1357) failed (9086)
15001 No volumes are able to be No volumes harvested. Make sure IBM StoredIQ
harvested in a job. For (15001) still has appropriate
instance, all of the mounts permissions to a volume.
fail due to a network issue. Verify that there is network
connectivity between the
data server and your
volume. Contact Customer
Support.
15002 Could not mount the volume. Error mounting volume Make sure that the data
Check permissions and <share><start-dir> on server still has appropriate
network settings. server <server-name>. permissions to a volume.
Reported <reason>. Verify that there is network
(15002) connectivity between the
data server and your
volume. Contact Customer
Support.
15021 Error saving harvest record. Failed to save Contact Customer Support.
HarvestRecord for qa1:auto- This message occurs due to
A (15021) a database error.
17501 Generic retention discovery Generic retention discovery Contact Customer Support.
failed in a catastrophic fatal failure: <17501>
manner.

104 IBM StoredIQ: Data Server Administration Guide


Table 38. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
17503 Generic retention discovery Error creating/loading Contact Customer Support.
creates volume sets volumeset for
associated with primary <server>:<share>
volumes. When that fails, IBM
StoredIQ sends this
message. This failure likely
occurred due to database
errors.
17505 Unable to query object count Unable to determine object Contact Customer Support.
for a discovered volume due count for <server>:<share>
to a database error.
17506 Generic retention discovery Error creating volume Contact Customer Support.
could not create discovered <server>:<share:>
volume.
18001 SMB connection fails. Windows Share Protocol Make sure IBM StoredIQ
Exception when connecting still has appropriate
to the server <server- permissions to a volume.
name> : <reason>. (18001) Verify that there is network
connectivity between the
data server and your
volume. Contact Customer
Support.
18002 The SMB volume mount Windows Share Protocol Verify the name of the
failed. Check the share name. Exception when connecting server and volume to make
to the share <share-name> sure that they are correct. If
on <server-name> : this message persists, then
<reason>. (18002) contact Customer Support.
18003 There is no volume manager. Windows Share Protocol Contact Customer Support.
Exception while initializing
the data object manager:
<reason>. (18003)
18006 Grazer volume crawl threw an Grazer._run : Unknown error Verify the user that mounted
exception. during walk. (18006) the specified volume has
permissions equivalent to
your current backup
solution. If this message
continues, contact
Customer Support.
18021 An unexpected error from the Unable to fetch trailing Check to ensure the
server prevented the harvest activity stream from NewsGator server has
to reach the end of the NewsGator volume. Will sufficient resources (disk
activity stream on the retry in next harvest. space, memory). It is likely
NewsGator data source that (18021) that this error is transient. If
is harvested. The next the error persists across
incremental harvest attempts multiple harvests, contact
to pick up from where the Customer Support.
current harvest was
interrupted.

Audits and logs 105


Table 38. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
18018 Start directory has escape Cannot graze the volume, Consider turning off escape
characters, and the data root directory Nunez has character checking.
server is configured to skip escape characters (18018)
them.
19001 An exception occurred during Interrogator._init Contact Customer Support.
interrogator initialization. __exception: <reason>.
(19001)
19002 An unknown exception Interrogator.__ Contact Customer Support.
occurred during interrogator
initialization. init __exception:
unknown. (19002)

19003 An exception occurred during [Link] Contact Customer Support.


interrogator processing. exception (<volumeid>,
<epoch>): <reason>.
(19003)

19004 An unknown exception [Link] Contact Customer Support.


occurred during interrogator exception (<volumeid>,
processing. <epoch>). (19004)

19005 An exception occurred during Viewer.__init__: Exception - Contact Customer Support.


viewer initialization. <reason>. (19005)
19006 An unknown exception Viewer.__init__: Unknown Contact Customer Support.
occurred during viewer exception. (19006)
initialization.
33003 Could not mount the volume. Unable to mount the Verify whether user name
Check permissions and volume: <error reason> and password that are used
network settings. (33003) for mounting the volume are
accurate. Check the user
data object for appropriate
permissions to the volume.
Make sure that the volume
is accessible from one of the
built-in protocols (NFS,
Windows Share, or
Exchange). Verify that the
network is properly
configured for the appliance
to reach the volume. Verify
that the appliance has
appropriate DNS settings to
resolve the server name.
33004 Volume could not be Unmounting volume failed Restart the data server. If
unmounted. from mount point : <mount the problem persists, then
point>. (33004) contact Customer Support.

106 IBM StoredIQ: Data Server Administration Guide


Table 38. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
33005 Data server was unable to Unable to create Restart the data server. If
create a local mounting point mount_point using the problem persists, then
for the volume. [Link]-Safe- contact Customer Support.
Makedirs(). (33005)
33010 Failed to make SMB Mounting Windows Share Verify user name and
connection to Windows Share volume failed with the password that is used for
server. error : <system error mounting the volume are
message>. (33010) accurate. Check the user
data object for appropriate
permissions to the volume.
Make sure that the volume
is accessible from one of the
built-in protocols (Windows
Share). Verify that the
network is properly
configured for the data
server to reach the volume.
Verify that the data server
has appropriate DNS
settings to resolve the
server name.
33011 Internal error. Problem Unable to open /proc/ Restart the data server. If
accessing local /proc/mounts mounts. Cannot test if the problem persists, then
volume was already contact Customer Support.
mounted. (33011)
33012 Database problems when a An exception occurred while Contact Customer Support.
volume was deleted. working with HARVESTS_
TABLE in Volume._delete().
(33012)
33013 No volume set was found for Unable to load volume set Contact Customer Support.
the volumes set name. by its name. (33013)
33014 System could not determine An error occurred while Contact Customer Support.
when this volume was last performing the last_harvest
harvested. operation. (33014)
33018 An error occurred mounting Mounting Exchange Server Verify user name and
the Exchange share. failed : <reason>. (33018) password that is used for
mounting the share are
accurate. Check for
appropriate permissions to
the share. Make sure that
the share is accessible.
Verify that the network is
properly configured for the
data server to reach the
share. Verify that the data
server has appropriate DNS
settings to resolve the
server name.

Audits and logs 107


Table 38. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
33027 The attempt to connect and Mounting IBM FileNet Ensure the connectivity,
authenticate to the IBM volume failed : <reason>. credentials, and
FileNet server failed. (33027) permissions to the FileNet
volume and try again.
33029 Failed to create the volume Exceeded maximum number Contact Customer Support.
as the number of active of volume partitions
volume partitions exceeds (33029).
the limit of 500.
34002 Could not complete the copy Copy Action aborted as the Verify that there is space
action because the target target disk has run out of available on your policy
disk was full. space (34002) destination and try again.
34009 Could not complete the move Move Action aborted as the Verify that there is space
action due to full target disk. target disk has run out of available on your policy
space.(34009) destination, then run
another harvest before you
run the policy. When the
harvest completes, try
running the policy again.
34015 The policy audit could not be Error Deleting Policy Audit: Contact Customer Support.
deleted for some reason. <error message> (34016)
34030 Discovery export policy is Production Run action Create sufficient space on
started since it detected the aborted because the target target disk and run
target disk is full. disk has run out of space. discovery export policy
(34030) again.
34034 The target volume for the Copy objects failed, unable Ensure the connectivity,
policy could not be mounted. to mount volume: login credentials, and
The policy is started. [Link]. permissions to the target
COM:SHARE. (34034) volume for the policy and try
again.
41004 The job is ended abnormally. <job-name> ended Try to run the job again. If it
unexpectedly. (41004) fails again, contact
Customer Support.
41007 Job failed. [Job name] has failed Look at previous messages
(41007). to see why it failed and refer
to that message ID to
pinpoint the error. Contact
Customer Support.
42001 The copy action could not run Copy data objects did not Contact Customer Support.
because of parameter errors. run. Errors occurred:<error-
description>. (42001)
42002 The copy action was unable Copy data objects failed, Check permissions on the
to create a target directory. unable to create target target. Make sure the
dir:<target-directory- permissions that are
name>. (42002) configured to mount the
target volume have write
access to the volume.

108 IBM StoredIQ: Data Server Administration Guide


Table 38. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
42004 An unexpected error Copy data objects Contact Customer Support.
occurred. terminated abnormally.
(42004)
42006 The move action could not Move data objects did not Contact Customer Support.
run because of parameter run. Errors occurred:<error-
errors. description>. (42006)
42007 The move action was unable Move data objects failed, Check permissions on the
to create a target directory. unable to create target target. Make sure the
dir:<target-directory- permissions that are
name>. (42007) configured to mount the
target volume have write
access to the volume.
42009 An unexpected error Move data objects Contact Customer Support.
occurred. terminated abnormally.
(42009)
42017 An unexpected error Delete data objects Contact Customer Support.
occurred. terminated abnormally.
(42017)
42025 The policy action could not Policy cannot execute. Contact Customer Support.
run because of parameter Attribute verification failed.
errors. (42025)
42027 An unexpected error Policy terminated Contact Customer Support.
occurred. abnormally. (42027)
42050 The data synchronizer could Content Data Synchronizer Contact Customer Support.
not run because of an synchronization of <server-
unexpected error. name>: <volume-name>
failed fatally.
42059 Invalid set of parameters that Production Run on objects Contact Customer Support.
are passed to discovery did not run. Errors occurred:
export policy. The following parameters
are missing: action_limit.
(42059)
42060 Discovery export policy failed Production Run on objects Verify that the discovery
to create target directory for (Copying native objects) export volume has write
the export. failed, unable to create permission and re-execute
target dir: production/10. policy.
(42060)
42062 Discovery export policy was Production Run on objects Contact Customer Support.
ended abnormally. (Copying native objects)
terminated abnormally.
(42062)
42088 The full-text optimization Full-text optimization failed Contact Customer Support.
process failed; however, the on volume <volume-name>
index is most likely still (42088)
usable for queries.

Audits and logs 109


Table 38. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
45802 A full-text index is already Time allocated to gain Contact Customer Support.
being modified. exclusive access to in-
memory index for volume=
1357 has expired (45802)
45803 The index for the specified Index '/deepfs/full-text/ No user intervention is
volume does not exist. This volume_index/ required.
message can occur under volume_1357' not found.
normal conditions. (45803)
45804 Programming error. A Transaction of client: Contact Customer Support.
transaction was never [Link].com_ FINDEX_
initiated or was closed early. QUEUE_ 1357_117251522
_3_ 2 is not the writer
(45804)
45805 The query is not started or Query ID: 123 does not exist No user intervention is
expired. The former is a (45805) required.
programming error. The latter
is normal.
45806 The query expression is Failed to parse 'dog pre\3 Revise your full-text query.
invalid or not supported. bar' (45806)
45807 Programming error. A Client: [Link].com_ Contact Customer Support.
transaction was already FINDEX_QUEUE
started for the client. _1357_1172515 222_3_2
is already active (45807)
45808 A transaction was never No transaction for client: No user intervention is
started or expired. [Link].com_ FINDEX_ required. The system
QUEUE_1357_ handles this condition
1172515222_3_2 (45808) internally.
45810 Programming error. Invalid volumeId. Expected: Contact Customer Support.
1357 Received:2468
(45810)
45812 A File I/O error occurred Failed to write disk (45812). Try your query again.
while the system was Contact Customer Support
accessing index data. for more assistance if
necessary.
45814 The query expression is too Query: 'a* b* c* d* e*' is too Refine your full-text query.
long. complex (45814)
45815 The file that is being indexed Java heap exhausted while Check the skipped file list in
is too large or the query indexing node with ID: the audit log for files that
expression is too complex. '10f4179cd5ff22f 2a6b failed to load due to their
The engine temporarily ran 79a1bc3aef247 fd94ccff' sizes. Revise your query
out of memory. (45815) expression and try again.
46023 Tar command failed while it Failed to back up full-text Check disk space and
persists full-text data to data for server:share. permissions.
Windows Share or NFS share. Reason: <reason>. (46023)

110 IBM StoredIQ: Data Server Administration Guide


Table 38. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
46024 Unhandled unrecoverable Exception <exception> Contact Customer Support.
exception while persisting while backing up fulltext
full-text data into a .tgz file. data for server:share
(46024)
46025 Was not able to delete Failed to unlink incomplete Check permissions.
partial .tgz file after a failed backup image. Reason:
full-text backup. <reason>. (46025)
47002 Synchronization failed on a Synchronization failed for Contact Customer Support.
query. query '<query-name>' on
volume '<server-and-
volume> (47002)
47101 An error occurred during the Cannot process full-text Restart services and contact
query of a full-text expression (Failed to read Customer Support.
expression. from disk (45812) (47101)
47203 No more database Database connections Contact Customer Support.
connections are available. exhausted (512/511)
(47203)
47207 User is running out of disk Disk usage exceeds Contact Customer Support.
space. threshold. (%d) In rare cases, this message
can indicate a program error
leaking disk space. In most
cases, however, disk space
is almost full, and more
storage is required.
47212 Interrogator failed while the Harvester 1 Does not exist. If the problem persists (that
system processed a file. The Action taken : restart. is, the system fails on the
current file is missing from (47212) same file or type of files),
the volume cluster. contact Customer Support.
47214 SNMP notification sender is Unable to resolve host name Check spelling and DNS
unable to resolve the trap [Link] [Link] setup.
host name. (47214)
50011 The DDL/DML files that are Database version control Contact Customer Support.
required for the database SQL file not found. (50011)
versioning were not found in
the expected location on the
data server.
50018 Indicates that the pre- Database restore is Contact Customer Support.
upgrade database restoration unsuccessful. Contact
failed, which was attempted Customer Support. (50018)
as a result of a database
upgrade failure.
50020 Indicates that the current Versions do not match! Contact Customer Support.
database requirements do Expected current database
not meet those requirements version: <dbversion>.
that are specified for the (50020)
upgrade and cannot proceed
with the upgrade.

Audits and logs 111


Table 38. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
50021 Indicates that the full Database backup failed. Contact Customer Support.
database backup failed when (50021)
the system attempts a data-
object level database backup.
61003 Discovery export policy failed Production policy failed to
to mount volume. mount volume. Aborting.
(61003)
61005 The discovery export load file Production load file Contact Customer Support.
generation fails generation failed. Load files
unexpectedly. The load files may be produced, but post-
can be produced correctly, processing may be
but post-processing actions incomplete. (61005)
like updating audit trails and
generating report files might
not complete.
61006 The discovery export load file Production load file Free up space on the target
generation was interrupted generation interrupted. disk, void the discovery
because the target disk is full. Target disk full. (61006) export run and run the
policy again.
68001 The gateway and data server Gateway connection failed Update your data server to
must be on the same version due to unsupported data the same build number as
to connect. server version. the gateway and restart
services. If your encounter
issues, contact Customer
Support.
68003 The data server failed to The data-server connection Contact Customer Support.
connect to the gateway over to the gateway cannot be
an extended period. established.
8002 The system failed to open a Failed to connect to the The "maximum database
connection to the database. database (80002) connections" configuration
parameter of the database
engine might need to be
increased. Contact
Customer Support.

INFO event log messages


The following table contains a complete listing of all INFO event-log messages.

Table 39. INFO event log messages


Event Required customer
number Reason Sample message action

9001 No conditions were added to Harvester: Query <query name> Add conditions to the
a query. cannot be inferred because no specified query.
condition for it has been defined
(9001).

112 IBM StoredIQ: Data Server Administration Guide


Table 39. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

9002 One or more conditions in a Harvester: Query < query name> Verify that regular
query were incorrect. cannot be inferred because of expressions are
regular expression or other properly formed.
condition error (9002).

9003 Volume Harvest is complete Volume statistics computation No user intervention is


and explorers are being started (9003). required.
calculated.

9004 Explorer calculations are Volume statistics computation No user intervention is


complete. completed (9004). required.

9005 Query membership Query inference will be done in No user intervention is


calculations started. <number> steps (9005). required.

9006 Query membership Query inference step <number> No user intervention is


calculations progress done (9006). required.
information.

9007 Query membership Query inference completed (9007). No user intervention is


calculations completed. required.

9012 Indicates the end of Dump of Volume cache(s) No user intervention is


dumping the content of the completed (9012). required.
volume cache.

9013 Indicates the beginning of Postprocessing for volume No user intervention is


the load process. 'Company Data Server:/mnt/demo- required.
A' started (9013).

9067 Indicates load progress. System metadata and tagged values No user intervention is
were successfully loaded for volume required.
'server:volume' (9067).

9069 Indicates load progress. Volume 'data server: /mnt/demo-A': No user intervention is
System metadata, tagged values required.
and full-text index were
successfully loaded (9069).

9084 The volume export finished. Exporting volume 'data server:/mnt/ No user intervention is
demo-A' (1357) completed (9084) required.

9087 The volume import finished. Importing volume 'dataserver:/mnt/ No user intervention is
demo-A' (1357) completed (9087) required.

9091 The load process was ended Load aborted due to user request No user intervention is
by the user. (9091). required.

15008 The volume load step was Post processing skipped for volume No user intervention is
skipped, per user request. <server>:<vo-lume>. (15008) required.

Audits and logs 113


Table 39. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

15009 The volume load step was Harvest skipped for volume No user intervention is
run but the harvest step was <server>:<vol-ume>. (15009) required.
skipped, per user request.

15012 The policy that ran on the Volume <volume> on server No user intervention is
volume is complete and the <server> is free now. Proceeding required.
volume load can now with load. (15012)
proceed.

15013 The configured time limit on Harvest time limit reached for No user intervention is
a harvest was reached. server:share. Ending harvest now. required.
(15013)

15014 The configured object count Object count limit reached for No user intervention is
limit on a harvest was server:share. Ending harvest now. required.
reached. (15014)

15017 checkbox is selected for Deferring post processing for No user intervention is
nightly load job. volume server:vol (15017) required.

15018 Harvest size or time limit is Harvest limit reached on No user intervention is
reached. server:volume. Synthetic deletes required.
will not be computed. (15018)

15019 User stops harvest process. Harvest stopped by user while No user intervention is
processing volume dpfsvr:vol1. Rest required.
of volumes will be skipped. (15019)

15020 The harvest vocabulary Vocabulary for dpfsvr:jhaide-A has Full harvest must be run
changed. Full harvest must changed. A full harvest is instead of an
run instead of incremental. recommended (15020). incremental harvest.

15022 The user is trying to run an Permission-only harvest: permission No action is needed as
ACL-only harvest on a checks not supported for the volume is skipped.
volume that is not a <server>:<share>
Windows Share or
SharePoint volume.

15023 The user is trying to run an Permission-only harvest: volume No action is needed as
ACL-only harvest on a <server>:<share> has no associated the volume is skipped.
volume that was assigned a user list.
user list.

17507 Limit (time or object count) Retention discovery limit reached Contact Customer
reached for generic retention for <server>:<share> Support.
discovery.

114 IBM StoredIQ: Data Server Administration Guide


Table 39. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

17508 Generic retention discovery No new items discovered. Post- No user intervention is
found no new items for this processing skipped for volume required unless the
master volume. <server>:<share> user is certain that new
items must be
discovered.

17509 Generic retention discovery Created new discovered volume No user intervention is
created a new volume. <server>:<share> in volume set required.
<autodis-covered volume set
name>.

18004 Job was stopped. Walker._process File: Grazer No user intervention is


Stopped. (18004) required.

18005 Grazer queue was closed. Walker._processFile: Grazer Closed. No user intervention is
(18005) required.

18016 Displays the list of top-level Choosing top-level directories: No user intervention is
directories that are selected <directories> (18016) required.
by matching the start
directory regular expression.
Displays at the beginning of
a harvest.

34001 Marks current progress of a <volume>: <count> data objects No user intervention is
copy action. processed by copy action. (34001) required.

34004 Marks current progress of a <volume>: <count> data objects No user intervention is
delete action. processed by delete action. (34004) required.

34008 Marks current progress of a <volume>: <count> data objects No user intervention is
move action. processed by move action. (34008) required.

34014 A policy audit was deleted. Deleting Policy Audit # <audit id> No user intervention is
<policy name> <start time> (34014) required.

34015 A policy audit was deleted. Deleted Policy Audit # <audit id> No user intervention is
<policy name> <start time> (34015) required.

34031 Progress update of the Winserver:top share : 30000 data No user intervention is
discovery export policy, objects processed by production required.
every 10000 objects action. (34031)
processed.

41001 A job was started either <jobname> started. (41001) No user intervention is
manually or was scheduled. required.

41002 The user stopped a job that <jobname> stopped at user request No user intervention is
was running. (41002) required.

41003 A job is completed normally <jobname> completed. (41003) No user intervention is


with or without success. required.

Audits and logs 115


Table 39. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

41006 Rebooting or restarting Service shutdown. Stopping Rerun jobs after restart
services on the controller or outstanding jobs. (41006) if you want the jobs to
compute node causes all complete.
jobs to stop.

41008 Database compactor Database compactor was not run Set the database
(vacuum) job cannot run because other jobs are active compactor's job
while there is database (41008). schedule so that it does
activity. not conflict with long-
running jobs.

42005 The action completed or was Copy complete: <number> data No user intervention is
ended. Shows results of objects copied,<number> collisions required.
copy action. found. (42005)

42010 The action completed or was Move complete: <number> data No user intervention is
ended. Shows results of objects moved,<number > collisions required.
move action. found.
(42010)

42018 The action completed or was Copy data objects complete: No user intervention is
ended. Shows results of <number> data objects required.
deleted action. copied,<number> collisions found.
(42018)

42024 The synchronizer was Content Data Synchronizer No user intervention is


completed normally. complete. (42024) required.

42028 The action completed or was Policy completed (42028). No user intervention is
ended. Shows results of required.
policy action.

42032 The action completed or was <report name> completed (42032). No user intervention is
ended. Shows results of required.
report action.

42033 The synchronizer started Content Data Synchronizer started. No user intervention is
automatically or manually (42033) required.
with the GUI button.

42048 Reports that the Content Data Synchronizer skipping No user intervention is
synchronizer is skipping a <server-name>:<volume-name> as required.
volume if synchronization is it does not need synchronization.
determined not to be 42048)
required.

42049 Reports that the Content Data Synchronizer starting No user intervention is
synchronizer started synchronization for volume <server- required.
synchronization of a volume. name>:<volume-name>

116 IBM StoredIQ: Data Server Administration Guide


Table 39. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

42053 The policy that was waiting Proceeding with execution of No user intervention is
for participant volumes to be <policy-name>. required.
loaded before it continues, is
now starting.

42063 Report on completion of Production Run on objects (Copying No user intervention is


discovery export policy native objects) completed: 2003 required.
execution phase. data objects copied, 25 duplicates
found. (42063)

42065 A discovery export policy Proceeding with execution of No user intervention is


that was held up for want of 'Production case One'. (42065) required.
resources, is now done
waiting, and begins
execution.

42066 A new discovery export run New run number 10 started for Note the new run
started. production Production Case 23221. number to tie the
(42066) current run with the
corresponding audit
trail.

42067 Discovery export policy is Production Run producing Audit No user intervention is
preparing the audit trail in Trail XML. (42067) required.
XML format. It might take a
few minutes.

42074 A query or tag was replicated Successfully sent query 'Custodian: No user intervention is
to a member data server Joe' to member data server San required.
successfully. Jose Office (42074)

46001 The backup process began. Backup Process Started. (46001) No user intervention is
Any selected backups in the required.
system configuration screen
are run if necessary.

46002 The backup process did not Backup Process Failed: <error- Check your backup
complete all its tasks description>. (46002) volume.
successfully. One or more
backup types did not occur.

46003 The backup process Backup Process Finished. (46003) No user intervention is
completed attempting all the required.
necessary tasks
successfully. Any parts of
the overall process add their
own log entries.

Audits and logs 117


Table 39. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

46004 The Application Data Application Data backup failed. Check your backup
backup, as part of the overall (46004) volume. Look at the
backup process, needed to setup for the
run but did not succeed. Application Data
backup. If backups
continue to fail, contact
Customer Support.

46005 The Application Data Application Data backup finished. No user intervention is
backup, as part of the overall (46005) required.
backup process, needed to
run and succeeded.

46006 The Application Data Application Data backup not No user intervention is
backup, as part of the overall configured, skipped. (46006) required.
backup process, was not
configured.

46007 The Harvested Volume Data Harvested Volume Data backup Check your backup
backup, as part of the overall failed. (46007) volume. Look at the
backup process, needed to setup for the Harvested
run but did not succeed. Volume Data backup. If
backups continue to
fail, contact Customer
Support.

46008 The Harvested Volume Data Harvested Volume Data backup No user intervention is
backup, as part of the overall finished. (46008) required.
backup process, needed to
run and succeeded.

46009 The Harvested Volume Data Harvested Volume Data backup not No user intervention is
backup, as part of the overall configured, skipped. (46009) required.
backup process, was not
configured.

46010 The System Configuration System Configuration backup failed. Check your backup
backup, as part of the overall (46010) volume. Look at the
backup process, needed to setup for the System
run but did not succeed. Configuration backup. If
backups continue to
fail, contact Customer
Support.

46011 The System Configuration System Configuration backup No user intervention is


backup, as part of the overall finished. (46011) required.
backup process, needed to
run and succeeded.

118 IBM StoredIQ: Data Server Administration Guide


Table 39. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

46012 The System Configuration System Configuration backup not No user intervention is
backup, as part of the overall configured, skipped. (46012) required.
backup process, was not
configured.

46013 The Audit Trail backup, as Policy Audit Trail backup failed. Check your backup
part of the overall backup (46013) volume. Look at the
process, needed to run but setup for the Audit Trail
did not succeed. backup. If back-ups
continue to fail, contact
IBM support.

46014 The Audit Trail backup, as Policy Audit Trail backup finished. No user intervention is
part of the overall backup (46014) required.
process, needed to run and
succeeded.

46015 The Audit Trail backup, as Policy Audit Trail backup not No user intervention is
part of the overall backup configured, skipped. (46015) required.
process was not configured.

46019 Volume cluster backup Indexed Data backup failed: Contact Customer
failed. <specific error> (46019) Support.

46020 Volume cluster backup Indexed Data backup finished. No user intervention is
finished. (46020) required.

46021 Volume is not configured for Indexed Data backup not No user intervention is
indexed data backups. configured, skipped. (46021) required.

46022 Full-text data was Successfully backed up full-text No user intervention is


successfully backed up. data for server:share (46022) required.

47213 Interrogator was Harvester 1 is now running. (47213) No user intervention is


successfully restarted. required.

60001 The user updates an object Query cities was updated by the No user intervention is
on the system. It includes administrator account (60001). required.
any object type on the data
server, including the
updating of volumes.

60002 The user creates an object. Query cities was created by the No user intervention is
It includes any object type administrator account (60002). required.
on the data server, including
the creation of volumes.

60003 The user deletes an object. Query cities was deleted by the No user intervention is
It includes any object type administrator account (60003). required.
on the data server, including
the deletion of volumes.

Audits and logs 119


Table 39. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

60004 The user publishes a full-text Query cities draft was published by No user intervention is
query set or a query. the administrator account (60004). required.

60005 The user tags an object. It Query tagging for cities class was No user intervention is
includes a published query, a started by the administrator account required.
draft query, or tag. (60005).

60006 A user restarted services on Application services restart for all No user intervention is
the data server. data servers was requested by the required.
administrator account (60006).

61001 Concordance discovery Preparing for upload of load file(s). No user intervention is
export is now preparing the (61001) required.
load files.

61002 Concordance discovery Load file(s) ready for upload. No user intervention is
export is ready to upload the (61002) required.
load files.

65000 The log file finished Log file download complete (65000) No user intervention is
downloading. required.

WARN event log messages


The following table contains a complete listing of all WARN event-log messages, reasons for occurrence,
sample messages, and any required customer action.

Table 40. WARN event log messages


Event
number Reason Sample message Customer action

1002 An Interrogator process failed Processing could not be Classify the document
because of an unknown error. completed on object, manually and Contact
The data object that was interrogator died : <data Customer Support.
processing is skipped. A new object name>. (1002)
process is created to replace
it.

1003 Interrogator child process did Interrogator terminated Try to readd the volume that
not properly get started. before accessing data is harvested. If that fails,
There might be problems to objects. (1003) contact Customer Support.
access the volume to be
harvested.

1004 Interrogator child process Processing was not Contact Customer Support.
was ended because it was no completed on object,
longer responding. The data interrogator killed : <data
object that was processing is object name>. (1004)
skipped. A new process is
created to replace it.

120 IBM StoredIQ: Data Server Administration Guide


Table 40. WARN event log messages (continued)
Event
number Reason Sample message Customer action

6001 A user email might not be Failed to send an email to Verify that your SMTP server
sent. The mail server settings user <email address>; check is configured correctly. Make
are incorrect. mail server configuration sure that the IP address that
settings (6001). is configured for the data
server can relay on the
configured SMTP server.

8001 The database needs to be The Database is approaching Run the database
vacuumed. an operational limit. Please maintenance task to vacuum
run the Database the database.
maintenance task using the
Console interface (8001)

9068 Tagged values were loaded, System metadata and tagged Contact Customer Support.
but full-text index loading values were loaded
failed. successfully for volume
'server:volume', but loading
the full-text index failed
(9068)

9070 Tagged values and full-text Loading system metadata, Contact Customer Support.
index loading failed. tagged values and the full-
text index failed for volume
'server:volume' (9070)

15003 The volume mount appeared Volume <volume name> on Contact Customer Support.
to succeed, but the test for server <server name> is not
mount failed. mounted. Skipping. (15003)

15004 A component cleanup failure [<component>] Cleanup Contact Customer Support.


on stop or completion. failure on stop. (15004)

15005 There was a component run [<component>] Run failure. Contact Customer Support.
failure. (15005)

15006 Cleanup failed for component [<component>] Cleanup Contact Customer Support.
after a run failure. failure on abort. (15006)

15007 A component that is timed Component [<component>] Try your action again. If this
out needs to be stopped. unresponsive; autostopping error continues, contact
triggered. (15007) Customer Support.

15010 The same volume cannot be Volume <volume-name> on No user intervention is


harvested in parallel. The server <server-name> is required. You might want to
harvest is skipped and the already being harvested. verify that the volume harvest
next one, if any are in queue, Skipping. (15010) is complete.
started.

Audits and logs 121


Table 40. WARN event log messages (continued)
Event
number Reason Sample message Customer action

15011 A volume cannot be Volume <volume-name> on No user intervention is


harvested if it is being used server <server-name> is required.
by another job. The harvest being used by another job.
continues when the job is Waiting before proceeding
complete. with load. (15011)

15015 Configured harvest time limit Time limit for harvest Reconfigure harvest time
is reached. reached. Skipping Volume v1 limit.
on server s1. 1 (15015)

15016 Configured harvest object Object count limit for harvest Reconfigure harvest data
count limit is reached. reached. Skipping Volume v1 object limit.
on server s1 (15016)

17008 Query that ran to discover Centera External Iterator : Contact Customer Support.
Centera items ended Centera Query terminated
unexpectedly. unexpectedly (<error
description>). (17008)

17011 Running discovery on the Pool Jpool appears to have Make sure that two jobs are
same pool in parallel is not another discovery running. not running at the same time
allowed. Skipping. (17011). that discovers the same pool.

17502 Generic retention discovery is Volume <server>:<share> No user intervention is


already running for this appears to have another required as the next step, if
master volume. discovery running. Skipping. any, within the job is run.

17504 Sent when a retention Volume <server>:<share> is Contact Customer Support.


discovery is run on any not supported for discovery.
volume other than a Windows Skipping.
Share retention volume.

18007 Directory listing or processing Walker._walktree: OSError - Make sure that the appliance
of data object failed in Grazer. <path><rea-son> (18007) still has appropriate
permissions to a volume.
Verify that there is network
connectivity between the
appliance and your volume.
Contact Customer Support.

18008 Unknown error occurred Walker._walktree: Unknown Contact Customer Support.


while processing data object exception - <path>. (18008)
or listing directory.

18009 Grazer timed out processing Walker._process File: Grazer Contact Customer Support.
an object. Timed Out. (18009)

18010 The skipdirs file is either not Unable to open skipdirs file: Contact Customer Support.
present or not readable by <filename>. Cannot skip
root. directories as configured.
(18010)

122 IBM StoredIQ: Data Server Administration Guide


Table 40. WARN event log messages (continued)
Event
number Reason Sample message Customer action

18011 An error occurred reading the Grazer._run: couldn't read Contact Customer Support.
known extensions list from extensions - <reason>.
the database. (18011)

18012 An unknown error occurred Grazer._run: couldn't read Contact Customer Support.
reading the known extensions extensions. (18012)
list from the database.

18015 NFS initialization warning that NIS Mapping not available. User name and group names
NIS is not available. (18015) might be inaccurate. Check
that your NIS server is
available and properly
configured in the data server.

18019 The checkpoint that is saved Unable to load checkpoint for If the message repeats in
from the last harvest of the NewsGator volume. A full subsequent harvests, contact
NewsGator data source failed harvest will be performed Customer Support.
to load. Instead of conducting instead. (18019)
an incremental harvest, a full
harvest is run.

18020 The checkpoint noted for the Unable to save checkpoint for If the message repeats in
current harvest of the NewsGator harvest of subsequent harvests, contact
NewsGator data source might volume. (18020) Customer Support.
not be saved. The next
incremental harvest of the
data source is not able to pick
up from this checkpoint.

33016 System might not unmount Windows Share Protocol Server administrators can see
this volume. Session teardown failed. that connections are left
(33016) hanging for a predefined time.
These connections will drop
off after they time out. No
user intervention required.

33017 System encountered an error An error occurred while Contact Customer Support.
while it tries to figure out retrieving the query instances
what query uses this volume. pointing to a volume. (33017)

33028 The tear-down operation of IBM FileNet tear-down None


the connection to a FileNet operation failed. (33028)
volume failed. Some
connections can be left open
on the FileNet server until
they time out.

34003 Skipped a copy data object Copy action error :- Target Verify that there is space
because disk full error. disk full, skipping copy : available on your policy
<source volume> to <target destination and try again.
volume>. (34003)

Audits and logs 123


Table 40. WARN event log messages (continued)
Event
number Reason Sample message Customer action

34010 Skipped a move data object Move action error :- Target Verify that there is space
because disk full error. disk full, skipping copy : available on your policy
<source volume> to <target destination. After verifying
volume>. (34010) that space is available, run
another harvest before you
run your policy. Upon harvest
completion, try running the
policy again.

34029 Discovery export policy Discovery export Run action Create sufficient space on
detects the target disk is full error: Target disk full, target disk and run discovery
and skips production of an skipping discovery export: export policy again.
object. share-1/saved/
[Link] to
production/10/
documents/1/0x0866e
5d6c898d9ffdbea 720b0
90a6f46d3058605 .txt.
(34029)

34032 The policy that was run has No volumes in scope for Check policy query and
no volumes in scope that is policy. Skipping policy scoping configuration, and re-
based on the configured execution. (34032) execute policy.
query and scoping. The policy
cannot be run.

34035 If the global hash setting for Copy objects : Target hash If target hashes need to be
the system is set to not will not be computed because computed for the policy audit
compute data object hash, no Hashing is disabled for trail, turn on the global hash
hash can be computed for the system. (34035) setting before you run the
target objects during a policy policy.
action.

34036 The policy has no source The policy has no source Confirm that the query used
volumes in scope, which volume(s) in scope. Wait for by the policy has one or more
means that the policy cannot the query to update before volumes in scope.
be run. executing the policy. (34036)

42003 The job in this action is Copy data objects stopped at No user intervention is
stopped by the user. user request. (42003) required.

42008 The job in this action is Move data objects stopped at No user intervention is
stopped by the user. user request. (42008) required.

42016 The job in this action is Delete data objects stopped No user intervention is
stopped by the user. at user request. (42016) required.

42026 The job in this action is Policy stopped at user No user intervention is
stopped by the user. request. (42026) required.

124 IBM StoredIQ: Data Server Administration Guide


Table 40. WARN event log messages (continued)
Event
number Reason Sample message Customer action

42035 When the job in this action is Set security for data objects No user intervention is
stopped by the user. stopped at user request. required.
(42035)

42051 Two instances of the same Policy <policy-name> is No user intervention is


policy cannot run at the same already running. Skipping. required.
time. (42051)

42052 Policies cannot be run until One or more volume(s) No user intervention is
after volumes are loaded, if needed by policy <policy- required.
those volumes are name> are being loaded.
participants to the policy by Waiting for the bulk load(s) to
virtue of being in the query. finish. (42052)

42061 Discovery export policy was Discovery export run on No user intervention is
stopped by user. objects (Copying native required.
objects) stopped at user
request. (42061)

42064 Discovery export policy A Discovery export run The discovery export policy
execution is delayed because related to policy 'Discovery execution is held up for
a conflicting discovery export export case One' is in required resources. Execution
run is in progress. progress. Waiting for it to must begin as soon as
finish. (42064) resource becomes available.

42068 Policy failed to set Copy objects warning, unable Policy might not be able to
appropriate permissions on to set permissions on target set appropriate permissions
the target directory. Objects directory: share-1/saved. on the objects it creates. If it
that are created from the (42068) is not acceptable, verify that
policy might not have target volume has
appropriate permissions set. appropriate write
permissions and re-execute.

42069 If the “Copy data objects Discovery export DAT_Export If the modified objects need
modified since last harvest” is configured to act on to be acted upon, either use a
option is selected for a members of containers, and discovery export action only
discovery export policy, it is cannot act on objects on the original file/email
valid only if the discovery modified after the last archive, or conduct an
export itself is defined to act harvest. Discovery export run incremental harvest on the
on the original file/email X will skip modified objects. source volumes.
archive, as opposed to their (42069)
members. If it is not true, the
warning tells the user that
modified objects are skipped.

46026 Volume is being harvested or Volume volume:share is in Rerun backup when volume is
policies are running against it. use. Unable to back up full- not in use.
If there are other full-text text index. Will retry later.
indexes to be backed up, the (46026)
system works on those
actions. Try this volume
again.

Audits and logs 125


Table 40. WARN event log messages (continued)
Event
number Reason Sample message Customer action

47201 Database connections are Database connections at No user intervention is


down to a normal level. normal level again (512/100) required.
(47201)

47202 The system is starting to run Database connections usage Contact Customer Support.
low on database connections. seems excessive (512/415)
This situation is abnormal. An (47202)
indication of process restarts
and connections are not
being cleared.

47215 Someone internally or SSHD: Failed password for Contact your local IT
externally is trying (and root from [Link] port manager. It might be either a
failing) to SSH into the data 57982. (47125) mistyped password by a
server. legitimate user or in the worst
case scenario, a genuine
break-in attempt.

61003 One of the load files cannot Failed to mount transaction Some of the load files will be
be uploaded because the cache dump '/deepfs/ missing after the discovery
compute node might not be postgres/pro-duction_cache'. export completes. These load
accessed to obtain. (61003) files are reproduced on a new
run. If problem persists
across runs, Contact
Customer Support.

61004 Warns the user that one of Transaction Cache Dump Run the discovery export
the transaction cache failed with error - Validation policy that saw the error
memory dump processes failed during creation of load again. If the error persists,
encountered an error. If a file. (61004) and you cannot find any
discovery export runs, it cluster/data server
means that the discovery configuration issues, contact
export fails to produce one of Customer Support.
the load files.
Note: If multiple memory
dumps fail, there is one
warning per failed memory
dump.

126 IBM StoredIQ: Data Server Administration Guide


Policy audits
Policy audits provide a detailed history of the policy. It includes type of action, date last run, start and end
dates with times, average speed, total data objects, and data object counts. They can be viewed by name,
volume, time, and by discovery export.

Policy audit by name

Table 41. Policy audit by name: Fields and descriptions


Policy audit by name
field Description
Policy name The policy name.
Policy status The policy's status.
Number of times The number of times that the policy was run.
executed
Most recent date The date on which the policy was last run.
executed

Policy audit by volume

Table 42. Policy audit by volume: Fields and descriptions


Policy audit by volume
field Description
Volume The name of the volume on which the policy was run.
Most recent date a The most recent date on which the policy was last run.
policy was executed
Number of policies The number of policies that were run.
executed

Policy audit by time

Table 43. Policy audit by time: Fields and descriptions


Policy audit by time
field Description
Policy name The policy name.
Policy status The status of the policy: Complete or Incomplete.
Start The time at which the policy's execution was started.
End The time at which the policy's execution was complete.
Success count The number of processed messages that are classified as a success.
Failure count The number of processed messages that are classified as a failure.
Warning count The number of processed messages that are classified as a warning.
Other count The number of processed messages that are classified as other.
Total data objects The total number of data objects.
Action type The type of policy that took place.

Audits and logs 127


Table 43. Policy audit by time: Fields and descriptions (continued)
Policy audit by time
field Description
Ag. actions/second The average number of actions per second.

Policy audit by discovery export

Table 44. Policy audit by discovery exports: Fields and descriptions


Policy audit by
discovery export field Description
Discovery export name The name of the discovery export.
Number runs The number of times the policy ran.
Most recent export The status of the most recent discovery export.
status
Most recent load file The status of the most recent load file.
status
Most recent date The date of the most recent policy execution.
executed

Discovery export runs by discovery export

Table 45. Discovery export runs by discovery export: Fields and descriptions
Discovery export runs
by discovery export
field Description
Discovery export run The name of the discovery export run.
Number of executions The number of times the run was started.
Success count The number of processed messages that are classified as a success.
Failure count The number of processed messages that are classified as a failure.
Warning count The number of processed messages that are classified as a warning.
Other count The number of processed messages that are classified as other.
Total data objects The total number of data objects.
Export status The status of the export: Complete or Incomplete.
Load file status The status of the load file: Complete or Incomplete.

Note: A warning in a policy audit trail is a success with the following conditions:
• If you copy an Exchange item such as re:, the re is copied, not the:. It generates a warning.
• The copied file is renamed.
• The file system to which you are copying does not accept characters in the file name.

Viewing policy audit details


Policy audits can be viewed by name, volume, time, or discovery export.
1. Go to Audit > Policies, and then click Name. The Policy audit by name page provides policy name and
status, the number of times it was run, and the time and date of the most recent execution.

128 IBM StoredIQ: Data Server Administration Guide


2. Click a policy name to open the Policy executions by time page.
3. Click a policy name to open the Policy execution results page.
Note: To view the list of data objects, click the [#] data objects link. To create a report, click Create
XML or Create PDF.
a) Click Volume to open the policy audit by volume page.
b) Click a volume link to go to the Policy audit by time page.
c) Click Time to see Audit by time page for the policy.
d) On the Policy audit by time page, click the policy name to open the Policy execution results page.
To view the list of data objects, click the [#] data objects link. To create a report, click Create XML
or Create PDF.
e) Click Discovery export.
f) On the Policy audit by discovery export page, click the discovery export name to open the
Discovery export runs by production page. The page details further information according to the
incremental runs of the policy.
g) Click a policy name to open the Policy executions by time page.
h) Click a policy name to open the Policy execution results page. To view the list of data objects, click
the [#] data objects link. To create a report, click Create XML or Create PDF.
As you review audit results through the pages, you can continue clicking through to review various
levels of information, from the volume and policy execution level down to the data objects. To view
more policy execution details, click the policy name in the execution summary page, which can be
accessed by any of the policy views. As you continue browsing, IBM StoredIQ provides more detailed
information such as:
• Source and destination settings
• Policy options: Details of the policy action. This section reflects the options that are selected when
you create the policy. Most attributes that appear depend upon the type of policy run and the options
available in the policy editor.
• Query (either IBM StoredIQ or user-defined)
• View metadata link: The view metadata page describes security details for source and destination
locations of the policy action.

Search audit feature


With the search audit feature, you can search audit trails by entering either by Policy Details, Execution
Details, or Data Object Details.

Policy details
Policy audits can be searched with any of these details.

Table 46. Policy audit details: Fields and descriptions


Policy detail field Description
Audit search by In this area, select search criteria, define their values, and then add them to the
policy details list to search across all audits.
Specify search In this area, specify the Policy name, the Policy state, and the Action type.
criteria
Audit search In the Find audits that match list, select either Any of the following or All of the
criteria following.

Execution details
Policy audits can be searched with any of these execution details.

Audits and logs 129


Table 47. Policy audit execution details: Fields and descriptions
Execution detail
field Description
Audit search by In this area, select search criteria, define their values, and then add them to the
execution detail list to search across all audits.
Specify search In this area, specify the Action type, Action status, Action start date, Action end
criteria date, Total count, Success count, Failure count, Warning count, Source
volume, Destination volume, or Query name.
Audit search In the Find audits that match list, select either Any of the following or All of the
criteria following.

Data object details


Policy audits can be searched with any of these data-object details.

Table 48. Policy audit data object details: Fields and descriptions
Data object details
field Description
Audit search by In this area, select search criteria, define their values, and then add them to the
data object details list to search across all audits.
Specify search In this area, specify the Source volume, Destination volume, Source object
criteria name, Destination object name, Source system path, Destination system path,
or Action result.
Audit search In the Find audits that match list, select either Any of the following or All of the
criteria following.

Saving results from an audit


You can save the results of policy executions into PDF and XML files. The information can be saved as PDF
and XML files. The exporting of information appears as a running job on the dashboard until completed.
1. Go to Audit > Policies.
2. In the Browse by options, click Time.
3. Click the policy name.
4. In the Results pane, click Data objects to see items that were responsive to the policy. To download
the material in .CSV, click CSV.
5. On the Policy execution results page, select Create PDF to generate a PDF or Create XML to generate
an XML file of the results.
6. Access the report through the inbox on the navigation page.

130 IBM StoredIQ: Data Server Administration Guide


Policy audit messages
A policy audit shows the number of data objects that were processed during the policy execution.
Processed data objects are divided into these categories: Success, Warnings, Failures, and Other
(discovery export policies only).

Table 49. Types of and reasons for policy audit messages


Audit message
type Reason
Success • Data object is a duplicate of [object name]
• Data object skipped but is loaded in load file. It applies to intermediate and files
archives produced during a discovery export policy.
• Data object is a duplicate produced in a previous run (discovery export only).

Warning • Set directory attributes


• Reset time stamps
• Set attributes
• Set time stamps
• Set security descriptor (Windows Share)
• Set access modes (Windows Share)
• Set owner information
• Set group information (NFS)
• Set security permissions
• Create a link after migration (Windows Share, NFS)
• Find template to create a shortcut (Windows Share)
• Extract text for the object (Discovery export policy)

Failure • Failed to create target directory structure


• Source does not exist
• Failed to find a new name for the incoming object
• Target is a directory
• File copy failed
• Cannot create target
• Error copying data to target
• Cannot copy due to network errors
• Cannot delete source after move
• Target disk is full
• Source equals target on a copy or move
• Insufficient permissions in general to conduct an action
• All modify actions failed
• File timed out waiting in the pipeline
• File under retention; cannot be deleted (retention server)
• Data object is a constituent of a container that already encountered failure (discovery
export policy)

Audits and logs 131


Table 49. Types of and reasons for policy audit messages (continued)
Audit message
type Reason
Other Data objects are categorized in the other category during a discovery export policy
when:
• A data object is a member that makes its container responsive.
• A data object is a non-responsive member of a container.

132 IBM StoredIQ: Data Server Administration Guide


Deploying SharePoint customized web services
This procedure highlights the basic steps that are required to deploy SharePoint custom web services.
This procedure applies only to SharePoint 2010, 2013, or 2016 web services.
1. Obtain the installation package.
This package is created for you by IBM StoredIQ. Depending on whether SharePoint 2010, 2013, or
2016 is used, the package is located in either /deepfs/downloads/webservice/sp2010/, /
deepfs/downloads/webservice/sp2013/, or /deepfs/downloads/webservice/sp2016/ on
the data server. Use a utility such as SCP to transfer the installation package.
2. Uninstall an existing instance of the web service.
To install an upgrade to a web service, any previous, existing instance must first be uninstalled. The
following steps must be completed on the SharePoint server.
a) Within IIS, click Sites and find the website that was created by the previous installation. Right-click
that website and remove it.
b) Within IIS, click Application Pools, find the web application that was created (it has the same
name as the website). Right-click that web application and remove it.
c) In Windows Explorer, go to the folder where the web service was deployed and delete all content
within this folder.
d) Reset IIS with the iisreset command.
3. Install the installation package.
4. Verify that the web service is hosted.
The following steps must be completed on the SharePoint server.
a) Within IIS, click Sites, and verify that you see the new site name that is listed along with the name
that is entered into the installer.
b) Expand Sites and verify that you can see the new site name that is listed along with the name that
is entered into the installer.
c) Select the newly created site and switch to the Content View, which is on the right pane.
d) An SVC file corresponds to the installed web service that is installed. Right-click the SVC file and
click Browse.
The web service is started in a browser, the address bar of which contains the HTTP location of the
web service such as [Link]
5. Configure admin knobs as described in “Configuring administration knobs” on page 134.

© Copyright IBM Corp. 2001, 2020 133


Configuring administration knobs
By configuring administration knobs you change settings that affect the IBM StoredIQ behavior.
The following administration knobs are provided:

Table 50. Admin knobs


Admin knob Description
The default value is 1.
cm8_missing_mime_type_error
If IBM Content Manager does not know about a
provided MIME type, an exception is generated. If
the user prefers that a generic MIME type
application/octet-stream is assigned to the
archived content instead of an exception, set this
knob to 1.

The default value is 0.


global_copy_ignore_target_vc
When set to 1, IBM StoredIQ does not
automatically harvest the destination volume when
it is creating copies.

Location of a custom web-service used to facilitate


sharepoint_custom_webservice_location
migration of time stamps and owner information in
the format port:servicelocation.
The default value is 0.
sharepoint_harvest_docs_only
When set to 1, IBM StoredIQ will harvest only
document libraries from SharePoint.

To configure an admin knob:


1. Using an SSH tool, log in to the data server VM as root.
2. Open a shell and run the following command:

psql -U dfuser -d dfdata

3. Use an UPDATE SQL statement to modify an admin knob record:

UPDATE adminknobs SET value='value' WHERE name='admin_knob_name';

Available admin knobs and any prerequisites are listed in Table 50 on page 134.
4. After your changes are complete, exit the psql utility.
5. Restart services by running this command:

service deepfiler restart

134 IBM StoredIQ: Data Server Administration Guide


Supported file types
The following section provides a comprehensive list of the file types that can be harvested and processed
by IBM StoredIQ, organized by name and by category. You can also view SharePoint attributes.

Supported file types by name


All file types by name that is supported by IBM StoredIQ are listed, including category, format, extension,
category, and version.

Table 51. Supported file types by name


Format Extension Category Version
Adobe Acrobat PDF graphic • 2.1
• 3.0-7.0
• Japanese

Adobe FrameMaker FMV graphic vector/raster through


Graphics 5.0
Adobe FrameMaker MIF word processing 3.0-6.0
Interchange Format
Adobe Illustrator graphic • Through 7.0
• 9.0

Adobe Photoshop PSD graphic 4.0


Ami Draw SDW graphic all
ANSI TXT text and markup 7- and 8-bit
ASCII TXT text and markup 7- and 8-bit
AutoCAD DWG CAD • 2.5-2.6
• 9.0-14.0
• 2002
• 2004
• 2005

AutoShade Rendering RND graphic 2.0


Binary Group 3 Fax graphic all
Bitmap BMP, RLE, ICO, CUR, graphic all
DIB, WARP
CALS Raster GP4 graphic Type I, II
Comma-Separated CSV spreadsheet
Values
Computer Graphics CGM graphic • ANSI
Metafile
• CALS
• NIST 3.0

© Copyright IBM Corp. 2001, 2020 135


Table 51. Supported file types by name (continued)
Format Extension Category Version
Corel Clipart CMX graphic 5-6
Corel Draw CDR graphic 3.x-8.x
Corel Draw (CDR with graphic 2.x-9.x
Tiff header)
Corel Presentations SHW presentation • Through 12.0
• X3

Corel WordPerfect WPD word processing • Through 12.0


Windows
• X3

DataEase Database 4.X


dBase Database Database Through 5.0
dBXL Database 1.3
DEC WPS PLUS DX word processing Through 4.0
DEC WPS PLUS WPL word processing Through 4.1
DisplayWrite (2 and 3) IP word processing all
DisplayWrite (4 and 5) word processing Through 2.0
DOS command COM system
executable
Dynamic link library files DLL system
EBCDIC text and markup all
ENABLE word processing • 3.0
• 4.0
• 4.5

ENABLE Database • 3.0


• 4.0
• 4.5

ENABLE Spreadsheet SSF spreadsheet • 3.0


• 4.0
• 4.5

Encapsulated Post- EPS graphic TIFF header


Script (raster)
Executable files EXE system
First Choice Database Through 3.0
First Choice word processing Through 3.0
First Choice spreadsheet Through 3.0
FoxBase Database 2.1

136 IBM StoredIQ: Data Server Administration Guide


Table 51. Supported file types by name (continued)
Format Extension Category Version
Framework Database 3.0
Framework word processing 3.0
Framework spreadsheet 3.0
GEM Bit Image IMG graphic all
Graphics Interchange GIF graphic all
Format
Graphics Environment GEM VDI graphic Bitmap and vector
Manager
Gzip GZ archive all
Haansoft Hangul HWP word processing • 1997
• 2002

Harvard Graphics (DOS) graphic • 2.x


• 3.x

Harvard Graphics graphic all


(Windows)
Hewlett-Packard HPGL graphic 2
Graphics Language
HTML HTM text and markup Through 3.0
IBM FFT text and markup all
IBM Graphics Data GDF graphic 1.0
Format
IBM Picture Interchange PIF graphic 1.0
Format
IBM Revisable Form text and markup all
Text
IBM Writing Assistant word processing 1.01
Initial Graphics IGES graphic 5.1
Exchange Spec
Java class files CLASS system
JPEG (not in TIFF JFIF graphic all
format)
JPEG JPEG graphic all
JustSystems Ichitaro JTD word processing • 5.0
• 6.0
• 8.0-13.0
• 2004

JustSystems Write word processing Through 3.0

Supported file types 137


Table 51. Supported file types by name (continued)
Format Extension Category Version
Kodak Flash Pix FPX graphic all
Kodak Photo CD PCD graphic 1.0
Legacy word processing Through 1.1
Legato Email Extender EMX Email
Lotus 1-2-3 WK4 spreadsheet Through 5.0
Lotus 1-2-3 (OS/2) spreadsheet Through 2.0
Lotus 1-2-3 Charts 123 spreadsheet Through 5.0
Lotus 1-2-3 for spreadsheet 1997-Millennium 9.6
SmartSuite
Lotus AMI Pro SAM word processing Through 3.1
Lotus Freelance PRZ presentation Through Millennium
Graphics
Lotus Freelance PRE presentation Through 2.0
Graphics (OS/2)
Lotus Manuscript word processing 2.0
Lotus Notes NSF Email
Lotus Pic PIC graphic all
Lotus Snapshot graphic all
Lotus Symphony spreadsheet • 1.0
• 1.1
• 2.0

Lotus Word Pro LWP word processing 1996-9.6


LZA Self Extracting archive all
Compress
LZH Compress archive all
Macintosh PICT1/2 PICT1/PICT1 graphic Bitmap only
MacPaint PNTG graphic NA
MacWrite II word processing 1.1
Macromedia Flash SWF presentation text only
MASS-11 word processing Through 8.0
Micrografx Designer DRW graphic Through 3.1
Micrografx Designer DSF graphic Win95, 6.0
Micrografx Draw DRW graphic Through 4.0

138 IBM StoredIQ: Data Server Administration Guide


Table 51. Supported file types by name (continued)
Format Extension Category Version
MPEG-1 Audio layer 3 MP3 multimedia • ID3 metadata only
• These files can be
harvested, but there is
no data in them that
can be used in tags.

MS Access MDB Database Through 2.0


MS Binder archive 7.0-1997
MS Excel XLS spreadsheet 2.2-2007
MS Excel Charts spreadsheet 2.x-7.0
MS Excel (Macintosh) XLS spreadsheet • 3.0-4.0
• 1998
• 2001
• 2004

MS Excel XML XLSX spreadsheet


MS MultiPlan spreadsheet 4.0
MS Outlook Express EML Email 1997-2003
MS Outlook Form OFT Email 1997-2003
Template
MS Outlook Message MSG Email all
MS Outlook Offline OST Email 1997-2003
Folder
MS Outlook Personal PST Email 1997-2007
Folder
MS PowerPoint PPT presentation 4.0-2004
(Macintosh)
MS PowerPoint PPT presentation 3.0-2007
(Windows)
MS PowerPoint XML PPTX presentation
MS Project MPP Database 1998-2003
MS Windows XML DOCX word processing
MS Word (Macintosh) DOC word processing • 3.0-4.0
• 1998
• 2001

MS Word (PC) DOC word processing Through 6.0


MS Word (Windows) DOC word processing Through 2007
MS WordPad word processing all
MS Works S30/S40 spreadsheet Through 2.0

Supported file types 139


Table 51. Supported file types by name (continued)
Format Extension Category Version
MS Works WPS word processing Through 4.0
MS Works (Macintosh) word processing Through 2.0
MS Works Database Database Through 2.0
(Macintosh)
MS Works Database (PC) Database Through 2.0
MS Works Database Database Through 4.0
(Windows)
MS Write word processing Through 3.0
Mosaic Twin spreadsheet 2.5
MultiMate 4.0 word processing Through 4.0
Navy DIF word processing all
Nota Bene word processing 3.0
Novell Perfect Works word processing 2.0
Novell Perfect Works spreadsheet 2.0
Novell Perfect Works graphic 2.0
(Draw)
Novell WordPerfect word processing Through 6.1
Novell WordPerfect word processing 1.02-3.0
(Macintosh)
Office Writer word processing 4.0-6.0
OpenOffice Calc SXC/ODS spreadsheet • 1.1
• 2.0

OpenOffice Draw graphic • 1.1


• 2.0

OpenOffice Impress SXI/SXP/ODP presentation • 1.1


• 2.0

OpenOffice Writer SXW/ODT word processing • 1.1


• 2.0

OS/2 PMMetafile MET graphic 3.0


Graphics
Paint Shop Pro 6 PSP graphic 5.0-6.0
Paradox Database (PC) Database Through 4.0
Paradox (Windows) Database Through 1.0
PC-File Letter word processing Through 5.0
PC-File+Letter word processing Through 3.0

140 IBM StoredIQ: Data Server Administration Guide


Table 51. Supported file types by name (continued)
Format Extension Category Version
PC PaintBrush PCX, DCX graphic all
PFS: Professional Plan spreadsheet 1.0
PFS: Write word processing A, B, C
Portable Bitmap Utilities PBM graphic all
Portable Greymap PGM graphic NA
Portable Network PNG graphic 1.0
Graphics
Portable Pixmap Utilities PPM graphic NA
PostScript File PS graphic level II
Professional Write word processing Through 2.1
Professional Write Plus word processing 1.0
Progressive JPEG graphic NA
Q &A (database) Database Through 2.0
Q & A (DOS) word processing 2.0
Q & A (Windows) word processing 2.0
Q & A Write word processing 3.0
Quattro Pro (DOS) spreadsheet Through 5.0
Quattro Pro (Windows) spreadsheet • Through 12.0
• X3

R:BASE 5000 Database Through 3.1


R:BASE (Personal) Database 1.0
R:BASE System V Database 1.0
RAR RAR archive
Reflex Database Database 2.0
Rich Text Format RTF text and markup all
SAMNA Word IV word processing
Smart Ware II Database 1.02
Smart Ware II word processing 1.02
Smart Ware II spreadsheet 1.02
Sprint word processing 1.0
StarOffice Calc SXC/ODS spreadsheet • 5.2
• 6.x
• 7.x
• 8.0

Supported file types 141


Table 51. Supported file types by name (continued)
Format Extension Category Version
StarOffice Draw graphic • 5.2
• 6.x
• 7.x
• 8.0

StarOffice Impress SXI/SXP/ODP presentation • 5.2


• 6.x
• 7.x
• 8.0

StarOffice Writer SXW/ODT word processing • 5.2


• 6.x
• 7.x
• 8.0

Sun Raster Image RS graphic NA


Supercalc Spreadsheet spreadsheet 4.0
Text Mail (MIME) various Email
Total Word word processing 1.2
Truevision Image TIFF graphic Through 6
Truevision Targa TGA graphic 2
Unicode Text TXT text and markup all
UNIX TAR (tape archive) TAR archive NA
UNIX Compressed Z archive NA
UUEncoding UUE archive NA
vCard word processing 2.1
Visio (preview) graphic 4
Visio 2003 graphic • 5
• 2000
• 2002

Volkswriter word processing Through 1.0


VP Planner 3D spreadsheet 1.0
WANG PC word processing Through 2.6
WBMP graphic NA
Windows Enhanced EMF graphic NA
Metafile
Windows Metafile WMF graphic NA
Winzip ZIP archive

142 IBM StoredIQ: Data Server Administration Guide


Table 51. Supported file types by name (continued)
Format Extension Category Version
WML text and markup 5.2
WordMARC word word processing Through Composer
processor
WordPerfect Graphics WPG, WPG2 graphic Through 2.0, 7. and 10
WordStar word processing Through 7.0
WordStar 2000 word processing Through 3.0
X Bitmap XBM graphic x10
X Dump XWD graphic x10
X Pixmap XPM graphic x10
XML (generic) XML text and markup
XyWrite XY4 word processing Through III Plus
Yahoo! IM Archive archive
ZIP ZIP archive PKWARE-2.04g

There is a potential for EBCDIC files to be typed incorrectly. It is especially true for raw text formats that
do not have embedded header information. In this case, IBM StoredIQ makes a best guess attempt at
typing the file.

Supported file types by category


All file types by category that is supported by IBM StoredIQ are listed, including category, format,
extension, and version.

Table 52. Supported archive file types by category


Format Extension Version
Gzip GZ all
LZA Self-Extracting Comparess all
LZH Compress all
MS Binder 7.0-1997
RAR RAR
Unix TAR (tape archive) TAR NA
Unix Compressed Z NA
UUEncoding UUE NA
Winzip Zip
Yahoo! IM Archive NA
ZIP ZIP PKWARE-2.04g

Supported file types 143


Table 53. Supported CAD file types by category
Format Extension Version
AutoCAD DWG • 2.5-2.6
• 9.0-14.0
• 2002
• 2004
• 2005

Table 54. Supported database file types by category


Format Extension Version
DataEase 4.x
dBase DataBase Through 5.0
dBXL 1.3
ENABLE • 3.0
• 4.0
• 4.5

First Choice Through 3.0


FoxBase 2.1
Framework 3.0
MS Access MDB Through 2.0
MS Project MPP Through 2.0
MS Works Database (Macintosh) 2.0
MS Works Database (PC) Through 2.0
MS Works Database (Windows) Through 4.0
Paradox Database (PC) Through 4.0
Paradox Database (Windows) Through 1.0
Q&A (database) Through 2.0
R:BASE 5000 Through 3.1
R:BASE (personal) 1.0
R:BASE System V 1.0
Reflex Database 2.0
Smart Ware II 1.02

Table 55. Supported email file types by category


Format Extension Version
Legato Email Extender EMX
Lotus Notes NSF

144 IBM StoredIQ: Data Server Administration Guide


Table 55. Supported email file types by category (continued)
Format Extension Version
MS Outlook Express EML 1997-2003
MS Outlook Form Template OFT 1997-2003
MS Outlook Message MSG all
MS Outlook Offline Folder OST 1997-2003
MS Outlook Personal Folder PST 1997-2007
Text Mail (MIME) various

Table 56. Supported graphic file types by category


Format Extension Version
Adobe Acrobat PDF • 2.1
• 3.0-7.0
• Japanese

Adobe Framemaker Graphics FMV vector/raster-5.0


Adobe Illustrator • Through 7.0
• 9.0

Adobe Photoshop PSD 4.0


Ami Draw SDW all
AutoShade Rendering RND 2.0
Binary Group 3 Fax all
Bitmap BMP, RLE, ICO, CUR, DIB, WARP all
CALS Raster GP4 Type I, II
Computer Graphics Metafile CGM • ANSI
• CALS
• NIST 3.0

Corel Clipart CMX 5-6


Corel Draw CDR 3.x-8.x
Corel Draw (CDR with TIFF 2.x-9.x
header
Encapsulated Post Script (raster) EPS TIFF header
GEM Bit Image IMG all
Graphics Interchange Format GIF all
Graphics Environment Manager GEM VDI • Bitmap
• vector

Harvard Graphics (DOS) • 2.x


• 3.x

Supported file types 145


Table 56. Supported graphic file types by category (continued)
Format Extension Version
Harvard Graphics (Windows) all
Hewlett-Packard Graphics HPGL 2
Language
IBM Graphics Data Format GDF 1.0
IBM Picture Interchange Format PIF 1.0
JPEG (not in TIFF format) JFIF all
JPEG JPEG all
Kodak Flash PIX FPX all
Kodak Photo CD PCD 1.0
Lotus Pic PIC all
Lotus Snapshot all
macintosh PICT1/2 PICT1/PICT2 Bitmap only
MacPaint PNTG NA
Micrografx Designer DRW Through 3.1
Micrografx Draw DRW Through 4.0
Novell Perfect Works (Draw) 2.0
OpenOffice Draw • 1.1
• 2.0

OZ/2 PM Metafile Graphics MET 3.0


Paint Shop Pro 6 PSP 5.0-6.0
PC Paintbrush PCX, DCX all
Portable Bitmap Utilities PBM all
Portable Network Graphics PNG 1.0
Portable Pixmap Utilities PPM NA
Postscript PS Level II
Progressive JPEG NA
StarOffice Draw • 5.2
• 6.x
• 7.x
• 8.0

Sun Raster Image RS NA


Truevision Image TIFF Through 6
Truevision Targa TGA 2
Visio 4

146 IBM StoredIQ: Data Server Administration Guide


Table 56. Supported graphic file types by category (continued)
Format Extension Version
Visio 2003 • 5
• 2000
• 2002

WBMP NA
Windows Enhanced Metafile EMF NA
Windows Metafile WMF NA
WordPerfect Graphics WPG, WPG2 • Through 2.0
• 7
• 10

X Bitmap XBM x10


XDump XWD x10
X Pixmap XPM x10

Table 57. Supported multimedia file types by category


Format Extension Version
MPEG-1 Audio Layer 3 MP3 ID3 metadata only
Note: These files can be
harvested, but there is no data in
them that can be used in tags.

Table 58. Supported presentation file types by category


Format Extension Version
Corel Presentations SHW • Through 12.0
• X3

Lotus Freelance Graphics PRZ Through Millennium


Lotus Freelance Graphics (OS/2) PRE Through 2.0
Macromedia Flash SWF Text only
MS PowerPoint (Macintosh) PPT 4.0-2004
MS PowerPoint (Windows) PPT 3.0-2007
MS PowerPoint XML PPTX
OpenOffice Impress SXI/SXP/ODP • 1.1
• 2.0

StarOffice Impress SXI/SXP/ODP • 5.2


• 6.x
• 7.x
• 8.0

Supported file types 147


Table 59. Supported spreadsheet file types by category
Format Extension Version
Comma-Separated Values CSV
ENABLE Spreadsheet SSF • 3.0
• 4.0
• 4.5

First Choice Through 3.0


Framework 3.0
Lotus 1-2-3 WK4 Through 5.0
Lotus 1-2-3 (OS/2) Through 2.0
Lotus 1-2-3 Charts 123 Through 5.0
Lotus 1-2-3 for SmartSuite 197-9.6
Lotus Symphony • 1.0
• 1.1
• 2.0

MS Excel XLS 2.2-2007


MS Excel Charts 2.x-7.0
MS Excel (Macintosh) XLS • 3.0-4.0
• 1998
• 2001
• 2004

MS Excel XML XLSX


MS MultiPlan 4.0
MS Works S30/S40 Through 2.0
Mosaic Twin 2.5
Novell Perfect Works 2.0
OpenOffice Calc SXC/ODS • 1.1
• 2.0

PFS: Professional Plan 1.0


Quattro Pro (DOS) Through 5.0
Quattro Pro (Windows) • Through 12.0
• X3

Smart Ware II 1.02

148 IBM StoredIQ: Data Server Administration Guide


Table 59. Supported spreadsheet file types by category (continued)
Format Extension Version
StarOffice Calc SXC/ODS • 5.2
• 6.x
• 7.x
• 8.0

Supercalc Spreadsheet 4.0


VP Planner 3D 1.0

Table 60. Supported system file types by category


Format Extension
Executable files .EXE
Dynamic link library files .DLL
Java class files .class
DOS command executables .COM

Table 61. Supported text and markup file types by category


Format Extension Version
ANSI .TXT 7-bit and 8-bit
ASCII .TXT 7-bit and 8-bit
EBCDIC all
HTML .HTM Through 3.0
IBM FFT all
IBM Revisable Form Text all
Rich Text Format RTF all
Unicode Text .TXT all
WML 5.2
XML .XML

There is a potential for EBCDIC files to be typed incorrectly. It is especially true for raw text formats that
do not have embedded header information. In this case, IBM StoredIQ makes a best guess attempt at
typing the file.

Table 62. Supported word-processing file types by category


Format Extension Version
Adobe FrameMaker Interchange MIF 3.0-6.0
Format
Corel WordPerfect Windows WPD • Through 12.0
• X3

DEC WPS PLUS DX Through 4.0

Supported file types 149


Table 62. Supported word-processing file types by category (continued)
Format Extension Version
DEC WPS PLUS WPL Through 4.1
Display Write (2 and 3) IP all
Display Write (4 and 5) Through 2.0
ENABLE • 3.0
• 4.0
• 4.5

First Choice Through 3.0


Framework 3.0
Haansoft Hangul HWP • 1997
• 2002

IBM Writing Assistant 1.01


JustSystems Ichitaro JTD • 5.0
• 6.0
• 8.0-13.0
• 2004

JustSystems Write Through 3.0


Legacy Through 1.1
Lotus AMI Pro SAM Through 3.1
Lotus Manuscript 2.0
Lotus Word Pro LWP 1996-9.6
MacWrite II 1.1
MASS-11 Through 8.0
MS Windows XML DOCX
MS Word (Macintosh) DOC • 3.0-4.0
• 1998
• 2001

MS Word (PC) DOC Through 6.0


MS Word (Windows) DOC Through 2007
MS WordPad all versions
MS Works WPS Through 4.0
MS Works (Macintosh) Through 2.0
MS Write Through 3.0
MultiMate 4.0 Through 4.0
Navy DIF all versions

150 IBM StoredIQ: Data Server Administration Guide


Table 62. Supported word-processing file types by category (continued)
Format Extension Version
Nota Bene 3.0
Novell Perfect Works 2.0
Novell WordPerfect Through 6.1
Novell WordPerfect (Macintosh) 1.02-3.0
Office Writer 4.0-6.0
OpenOffice Writer SXW/ODT • 1.1
• 2.0

PC-File Letter Through 5.0


PC-File + Letter Through 3.0
PFS Write • A
• B
• C

Professional Write Plus Through 2.1


Q&A (DOS) 2.0
Q&A (Windows) 2.0
Q&A Write 3.0
SAMNA Word IV
Smart Ware II 1.02
Sprint 1.0
StarOffice Writer SXW/ODT • 5.2
• 6.x
• 7.x
• 8.0

Total Word 1.2

SharePoint supported file types


The following section describes the various SharePoint data object types and their properties that are
currently supported by IBM StoredIQ.

Supported SharePoint object types


These types of SharePoint objects are supported:

Table 63. Supported SharePoint object types


Supported SharePoint object Supported SharePoint object Supported SharePoint object
types types types

• Blog posts and comments • Discussion board • Calendar

Supported file types 151


Table 63. Supported SharePoint object types (continued)
Supported SharePoint object Supported SharePoint object Supported SharePoint object
types types types

• Tasks • Project tasks • Contacts

• Wiki pages • Issue tracker • Announcements

• Survey • Links • Document libraries

• Picture libraries • Records center

Notes regarding SharePoint object types


• Calendar: Recurring calendar events are indexed as a single object in IBM StoredIQ. Each recurring
calendar event has multiple Event Date and End Date attribute values, one pair per recurrence. For
instance, if there is an event defined for American Independence Day and is set to recur yearly, it is
indexed with Event Dates 2010-07-04, 2011-07-04, 2012-07-04, and more.
• Survey: Only individual responses to a survey are indexed as system-level objects. Each response is a
user's feedback to all questions in the survey. Each question in the survey that was answered for a
response is indexed as an attribute of the response in the IBM StoredIQ index. The name of the
attribute is the string that forms the question while the value is the reply entered. Surveys have no full-
text indexable body, and they are always indexed with size=0.

Hash computation
The hash of a full-text indexed object is computed with the full-text indexable body of the object.
However, in the case of SharePoint list item objects (excluding documents and pictures), the full-text
indexable body might be empty or too simplistic. It means that you can easily obtain duplicate items
across otherwise two different objects. For this reason, other attributes are included in the hash
computation algorithm.
These attributes are included while the hash is computed for the SharePoint data objects, excluding
documents and pictures.

Table 64. Attribute summary


Attribute Types
Generic attributes • Title (SharePoint)
• Content type (SharePoint)
• Description (SharePoint)

Blog post attributes • Post category (SharePoint)

Wiki page attributes • Wiki page comment

Calendar event attributes • Event category (SharePoint)


• Event date (SharePoint)
• Event end date (SharePoint)
• Event location (SharePoint)

152 IBM StoredIQ: Data Server Administration Guide


Table 64. Attribute summary (continued)
Attribute Types
Task or project task attributes • Task start date (SharePoint)
• Task due date (SharePoint)
• Task that is assigned to (SharePoint)

Contact attributes • Contact full name(SharePoint)


• Contact email (SharePoint)
• Contact job title (SharePoint)
• Contact work address (SharePoint)
• Contact work phone (SharePoint)
• Contact home phone (SharePoint)
• Contact mobile phone (SharePoint)

Link attributes • Link URL (SharePoint)

Survey attributes • All survey questions and answers in the response


are included in the hash.

Supported file types 153


Supported server platforms and protocols
The following section lists the supported server platforms by volume type and the protocols for supported
systems.
Primary volume
A primary volume is storage knowledge workers access to create, read, update, and delete
unstructured content. Unstructured content is stored in standard formats such as office documents,
text files, system logs, application logs, email, and compressed archives that contain email,
documents, or enterprise social media content.
Retention volume
Generally, a retention volume is immutable storage that enforces retention and hold policies. While
data is under management, it cannot be modified or deleted, and knowledge workers typically do not
access retention storage directly. Specific applications move data to this storage to manage it. This
storage typically has its own special API/protocol, although NAS vendors implemented retention/hold
features that use standard CIFS/SMB or SMB2 and NFS protocols. The storage platform typically does
not implement a hierarchical namespace to store content, but instead relies on a globally unique
identifier as a handle to metadata and content. Applications are free to write metadata and binary
content in any internal format to satisfy their requirements.
Typically, IBM StoredIQ does not attempt to discover and manage data that is written by other
applications to retention volumes. Application-specific knowledge is often required to interpret
metadata and content. Retention storage is used by IBM StoredIQ to manage data on compliant
immutable storage for retention and holds. It does so in a way that does not interfere with knowledge
workers that create, update, and access content on primary volumes.
IBM StoredIQ preserves source metadata when data is written to a retention volume. The original
source metadata is important for governance and legal discovery (custodian, time stamps, and more)
to replicate the content and metadata from the retention volume when needed.
Export volume
Export volumes are unmanaged (not indexed) storage location where content is copied along with
metadata and audit detail in a format that can be imported by other applications. A common usage of
export volumes is to stage local documents to be imported into a legal review tool in a format such as
standard EDRM or a Concordance-compatible format.
System volume
System volumes are a storage location where files can be written to and read by IBM StoredIQ. It can
be used to export volume metadata that is contained in the index on a Data Server. Exported volume
data can be imported from a system volume to populate a volume index.

Supported platforms and protocols by IBM StoredIQ volume type

IBM StoredIQ IBM StoredIQ IBM StoredIQ


Platform/ primary retention IBM StoredIQ system
protocol volume volume export volume volume Notes
Box x Box volumes are
added through
IBM StoredIQ
Administrator.
CIFS/SMB or x x x x
SMB2
CMIS 1.0 x
Connections x

154 IBM StoredIQ: Data Server Administration Guide


IBM StoredIQ IBM StoredIQ IBM StoredIQ
Platform/ primary retention IBM StoredIQ system
protocol volume volume export volume volume Notes
EMC x Customer must
Documentum supply DFC files
to enable the
connector.
HDFS x
(HADOOP)
IBM Content x
Manager
IBM Domino x Email only.
This includes
IBM Verse®.

IBM FileNet x
Jive x
Microsoft x
Exchange
Microsoft x
SharePoint
NewsGator x
NFS x x x x
OneDrive for x
Business

Supported server platforms and protocols 155


IBM StoredIQ IBM StoredIQ IBM StoredIQ
Platform/ primary retention IBM StoredIQ system
protocol volume volume export volume volume Notes
OpenText x A copy of the
Livelink/ [Link] file
Content Server from a Livelink
API installation
must be
available in
the /usr/
local/IBM/IC
I/vendor
directory on
each data server
in your
deployment.
Usually, you can
find this file on
the Livelink
server in the
C:\OPENTEXT
\application
\WEB-INF\lib
directory.
However, the
path might be
different in your
Livelink
installation.
Salesforce x
Chatter

If IBM StoredIQ Desktop Data Collector is installed, Windows desktops can serve as primary volumes.
Supported actions for desktops are copy from, move from, export from, and delete. For more information,
see “Desktop collection” on page 94.

Supported operations and limitations by platform and protocol

Operation type: Operation type: Operation type:


Platform/protocol read write delete Notes
Box R W D Box volumes are
added through IBM
StoredIQ
Administrator.
CIFS/SMB or SMB2 R W D
Connections R

156 IBM StoredIQ: Data Server Administration Guide


Operation type: Operation type: Operation type:
Platform/protocol read write delete Notes
CMIS 1.0 R W The CMIS option
references a
solution that
involves different
products. As such,
no specific
supported version
is to cite, and it is
not called out in
the following
tables.
EMC Documentum R W D Only Current
documents in
standard cabinets
are indexed.
HDFS (HADOOP) R W D
IBM Content R W
Manager
IBM Domino R Email only. Email is
converted to
the .MSG format
for processing.
This includes IBM
Verse.

IBM FileNet R W
Jive R
Microsoft Exchange R Messages (email),
Contacts, Calendar
items, Notes,
Tasks, and
Documents

Supported server platforms and protocols 157


Operation type: Operation type: Operation type:
Platform/protocol read write delete Notes
Microsoft R W D All document
SharePoint versions are
optional.
SharePoint
2010/2013/2016
supported list
types: User
profiles, User
notes, Blog Post,
Blog Comment,
Discussion Post,
Discussion Reply,
Wiki Page,
Calendar, Task/
Project Task,
Contact, Issue
Tracker, Survey,
Link, and
Announcements.
Content of custom
lists is indexed
generically as text.
It is not modeled
specifically, like
standard list types.

NewsGator R No API support for


Poll responses.
NFS R W D
OneDrive for R W
Business
OpenText Livelink/ R
Content Server
Salesforce Chatter R

Supported CIFS/SMB or SMB2 server platform and protocol versions (read/write/delete)

CIFS/SMB or SMB2 server Notes


Windows XP, Vista, 7, 8, 10
Windows Server 2003, 2008, 2008 R2, 2012, 2016
Samba
Mac OS X, 10.7, 10.8, 10.9

Supported Connections server platform and protocol versions (read/write/delete)

Connections server Notes


Connections 5.0

158 IBM StoredIQ: Data Server Administration Guide


Connections server Notes
Connections 6.0

Supported EMC Documentum server platform and protocol versions (read/write)

EMC Documentum server Notes


Documentum 6.0, 6.5, 6.7 Customer must supply DFC files to enable the
connector.
Documentum with Retention Policy Services (RPS) Customer must supply DFC files to enable the
6.0, 6.5 connector.

Supported IBM FileNet server platform and protocol versions (read/write)

FileNet content services Notes


FileNet [Link]

Supported HDFS (HADOOP) server platform and protocol versions (read/write/delete)

HDFS service Notes


HADOOP v 2.7

Supported IBM Content Manager server platform and protocol versions (read/write)

IBM Content Manager 8.4.3 and later Notes


AIX 5L 5.3, 6.1, 7.1
DB2® (back-end database)

Red Hat® Enterprise Linux® 4.0, 5.0


SUSE Linux Enterprise Server 9, 10, 11

Oracle (back-end database)


Solaris 9, 10
Windows Server 2003, 2008, 2008 R2

Supported IBM Domino/Notes server platform and protocol versions (read)

IBM Domino/Notes server Notes


Domino/Notes 6.x, 7.x, 8.x, 9.x

Supported Jive server platform and protocol versions (read)

Jive service Notes


Jive 5.0.2

Supported Microsoft Exchange server platform and protocol versions (read)

Microsoft Exchange server Notes


Exchange 2003 WebDAV protocol

Supported server platforms and protocols 159


Microsoft Exchange server Notes
Exchange 2007, 2010, 2013, 2016, Online Exchange web service interface

Supported Microsoft SharePoint server platform and protocol versions (read/write/delete)

Microsoft SharePoint server Notes


SharePoint 2003 WebDAV protocol
SharePoint 2007, 2010, 2013, 2016, Online SharePoint web service interface

Supported NewsGator server platform and protocol versions (read)

NewsGator API Notes


NewsGator Social 2.1.1229 Installed on SharePoint 2010 or later.

Supported NFS server platform and protocol versions (read/write/delete)

NFSv3 Notes
Red Hat Enterprise Linux 5.x, 6.x
CentOS 5.x, 6.x
Mac OS X, 10.7, 10.8, 10.9

Supported OneDrive server platform and protocol versions (read/write)

OneDrive Notes
[Link] HTTP REST API

Supported OpenText Livelink Enterprise Server platform and protocol versions (read)

Livelink Enterprise Server Notes


OpenText Livelink Enterprise Server 9.7, 9.7.1 Connector is based on IBM Content Integrator (ICI)
version 8.6.
A copy of the [Link] file from a Livelink API
installation must be available in the /usr/
local/IBM/ICI/vendor directory on each data
server in your deployment. Usually, you can find
this file on the Livelink server in the C:\OPENTEXT
\application\WEB-INF\lib directory.
However, the path might be different in your
Livelink installation.

Supported OpenText Content Server platform and protocol versions (read)

Content Server Notes


Content Server 10.0.0 Connector is based on IBM Content Integrator (ICI)
version 8.6.

160 IBM StoredIQ: Data Server Administration Guide


Supported Salesforce Chatter server platform and protocol versions (read)

Salesforce Chatter service Notes


Salesforce Partner API v26.0

Supported server platforms and protocols 161


Event log messages
The following sections contain a complete listing of all ERROR, INFO, and WARN event-log messages that
appear in the Event Log of the IBM StoredIQ console.
• “ERROR event log messages” on page 162
• “INFO event log messages” on page 170
• “WARN event log messages” on page 178

ERROR event log messages


The following table contains a complete listing of all ERROR event-log messages, reasons for occurrence,
sample messages, and any required customer action.

Table 65. ERROR event log messages. This table lists all ERROR event-log messages.
Event
number Reason Sample message Required customer action
1001 Harvest was unable to open a Harvest could not allocate Log in to UTIL and restart
socket for listening to child listen port after <number> the application server.
process. attempts. Cannot kickstart Restart the data server.
interrogators. (1001) Contact customer support.
9083 Unexpected error while it is Exporting volume Contact Customer Support.
exporting a volume. 'dataserver:/mnt/demo-A'
(1357) has failed (9083)
9086 Unexpected error while it is Importing volume Contact Customer Support.
importing a volume 'dataserver:/mnt/demo-A'
(1357) failed (9086)
15001 No volumes are able to be No volumes harvested. Make sure IBM StoredIQ
harvested in a job. For (15001) still has appropriate
instance, all of the mounts permissions to a volume.
fail due to a network issue. Verify that there is network
connectivity between the
data server and your
volume. Contact Customer
Support.
15002 Could not mount the volume. Error mounting volume Make sure that the data
Check permissions and <share><start-dir> on server still has appropriate
network settings. server <server-name>. permissions to a volume.
Reported <reason>. Verify that there is network
(15002) connectivity between the
data server and your
volume. Contact Customer
Support.
15021 Error saving harvest record. Failed to save Contact Customer Support.
HarvestRecord for qa1:auto- This message occurs due to
A (15021) a database error.
17501 Generic retention discovery Generic retention discovery Contact Customer Support.
failed in a catastrophic fatal failure: <17501>
manner.

162 IBM StoredIQ: Data Server Administration Guide


Table 65. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
17503 Generic retention discovery Error creating/loading Contact Customer Support.
creates volume sets volumeset for
associated with primary <server>:<share>
volumes. When that fails, IBM
StoredIQ sends this
message. This failure likely
occurred due to database
errors.
17505 Unable to query object count Unable to determine object Contact Customer Support.
for a discovered volume due count for <server>:<share>
to a database error.
17506 Generic retention discovery Error creating volume Contact Customer Support.
could not create discovered <server>:<share:>
volume.
18001 SMB connection fails. Windows Share Protocol Make sure IBM StoredIQ
Exception when connecting still has appropriate
to the server <server- permissions to a volume.
name> : <reason>. (18001) Verify that there is network
connectivity between the
data server and your
volume. Contact Customer
Support.
18002 The SMB volume mount Windows Share Protocol Verify the name of the
failed. Check the share name. Exception when connecting server and volume to make
to the share <share-name> sure that they are correct. If
on <server-name> : this message persists, then
<reason>. (18002) contact Customer Support.
18003 There is no volume manager. Windows Share Protocol Contact Customer Support.
Exception while initializing
the data object manager:
<reason>. (18003)
18006 Grazer volume crawl threw an Grazer._run : Unknown error Verify the user that mounted
exception. during walk. (18006) the specified volume has
permissions equivalent to
your current backup
solution. If this message
continues, contact
Customer Support.
18021 An unexpected error from the Unable to fetch trailing Check to ensure the
server prevented the harvest activity stream from NewsGator server has
to reach the end of the NewsGator volume. Will sufficient resources (disk
activity stream on the retry in next harvest. space, memory). It is likely
NewsGator data source that (18021) that this error is transient. If
is harvested. The next the error persists across
incremental harvest attempts multiple harvests, contact
to pick up from where the Customer Support.
current harvest was
interrupted.

Event log messages 163


Table 65. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
18018 Start directory has escape Cannot graze the volume, Consider turning off escape
characters, and the data root directory Nunez has character checking.
server is configured to skip escape characters (18018)
them.
19001 An exception occurred during Interrogator._init Contact Customer Support.
interrogator initialization. __exception: <reason>.
(19001)
19002 An unknown exception Interrogator.__ Contact Customer Support.
occurred during interrogator
initialization. init __exception:
unknown. (19002)

19003 An exception occurred during [Link] Contact Customer Support.


interrogator processing. exception (<volumeid>,
<epoch>): <reason>.
(19003)

19004 An unknown exception [Link] Contact Customer Support.


occurred during interrogator exception (<volumeid>,
processing. <epoch>). (19004)

19005 An exception occurred during Viewer.__init__: Exception - Contact Customer Support.


viewer initialization. <reason>. (19005)
19006 An unknown exception Viewer.__init__: Unknown Contact Customer Support.
occurred during viewer exception. (19006)
initialization.
33003 Could not mount the volume. Unable to mount the Verify whether user name
Check permissions and volume: <error reason> and password that are used
network settings. (33003) for mounting the volume are
accurate. Check the user
data object for appropriate
permissions to the volume.
Make sure that the volume
is accessible from one of the
built-in protocols (NFS,
Windows Share, or
Exchange). Verify that the
network is properly
configured for the appliance
to reach the volume. Verify
that the appliance has
appropriate DNS settings to
resolve the server name.
33004 Volume could not be Unmounting volume failed Restart the data server. If
unmounted. from mount point : <mount the problem persists, then
point>. (33004) contact Customer Support.

164 IBM StoredIQ: Data Server Administration Guide


Table 65. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
33005 Data server was unable to Unable to create Restart the data server. If
create a local mounting point mount_point using the problem persists, then
for the volume. [Link]-Safe- contact Customer Support.
Makedirs(). (33005)
33010 Failed to make SMB Mounting Windows Share Verify user name and
connection to Windows Share volume failed with the password that is used for
server. error : <system error mounting the volume are
message>. (33010) accurate. Check the user
data object for appropriate
permissions to the volume.
Make sure that the volume
is accessible from one of the
built-in protocols (Windows
Share). Verify that the
network is properly
configured for the data
server to reach the volume.
Verify that the data server
has appropriate DNS
settings to resolve the
server name.
33011 Internal error. Problem Unable to open /proc/ Restart the data server. If
accessing local /proc/mounts mounts. Cannot test if the problem persists, then
volume was already contact Customer Support.
mounted. (33011)
33012 Database problems when a An exception occurred while Contact Customer Support.
volume was deleted. working with HARVESTS_
TABLE in Volume._delete().
(33012)
33013 No volume set was found for Unable to load volume set Contact Customer Support.
the volumes set name. by its name. (33013)
33014 System could not determine An error occurred while Contact Customer Support.
when this volume was last performing the last_harvest
harvested. operation. (33014)
33018 An error occurred mounting Mounting Exchange Server Verify user name and
the Exchange share. failed : <reason>. (33018) password that is used for
mounting the share are
accurate. Check for
appropriate permissions to
the share. Make sure that
the share is accessible.
Verify that the network is
properly configured for the
data server to reach the
share. Verify that the data
server has appropriate DNS
settings to resolve the
server name.

Event log messages 165


Table 65. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
33027 The attempt to connect and Mounting IBM FileNet Ensure the connectivity,
authenticate to the IBM volume failed : <reason>. credentials, and
FileNet server failed. (33027) permissions to the FileNet
volume and try again.
33029 Failed to create the volume Exceeded maximum number Contact Customer Support.
as the number of active of volume partitions
volume partitions exceeds (33029).
the limit of 500.
34002 Could not complete the copy Copy Action aborted as the Verify that there is space
action because the target target disk has run out of available on your policy
disk was full. space (34002) destination and try again.
34009 Could not complete the move Move Action aborted as the Verify that there is space
action due to full target disk. target disk has run out of available on your policy
space.(34009) destination, then run
another harvest before you
run the policy. When the
harvest completes, try
running the policy again.
34015 The policy audit could not be Error Deleting Policy Audit: Contact Customer Support.
deleted for some reason. <error message> (34016)
34030 Discovery export policy is Production Run action Create sufficient space on
started since it detected the aborted because the target target disk and run
target disk is full. disk has run out of space. discovery export policy
(34030) again.
34034 The target volume for the Copy objects failed, unable Ensure the connectivity,
policy could not be mounted. to mount volume: login credentials, and
The policy is started. [Link]. permissions to the target
COM:SHARE. (34034) volume for the policy and try
again.
41004 The job is ended abnormally. <job-name> ended Try to run the job again. If it
unexpectedly. (41004) fails again, contact
Customer Support.
41007 Job failed. [Job name] has failed Look at previous messages
(41007). to see why it failed and refer
to that message ID to
pinpoint the error. Contact
Customer Support.
42001 The copy action could not run Copy data objects did not Contact Customer Support.
because of parameter errors. run. Errors occurred:<error-
description>. (42001)
42002 The copy action was unable Copy data objects failed, Check permissions on the
to create a target directory. unable to create target target. Make sure the
dir:<target-directory- permissions that are
name>. (42002) configured to mount the
target volume have write
access to the volume.

166 IBM StoredIQ: Data Server Administration Guide


Table 65. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
42004 An unexpected error Copy data objects Contact Customer Support.
occurred. terminated abnormally.
(42004)
42006 The move action could not Move data objects did not Contact Customer Support.
run because of parameter run. Errors occurred:<error-
errors. description>. (42006)
42007 The move action was unable Move data objects failed, Check permissions on the
to create a target directory. unable to create target target. Make sure the
dir:<target-directory- permissions that are
name>. (42007) configured to mount the
target volume have write
access to the volume.
42009 An unexpected error Move data objects Contact Customer Support.
occurred. terminated abnormally.
(42009)
42017 An unexpected error Delete data objects Contact Customer Support.
occurred. terminated abnormally.
(42017)
42025 The policy action could not Policy cannot execute. Contact Customer Support.
run because of parameter Attribute verification failed.
errors. (42025)
42027 An unexpected error Policy terminated Contact Customer Support.
occurred. abnormally. (42027)
42050 The data synchronizer could Content Data Synchronizer Contact Customer Support.
not run because of an synchronization of <server-
unexpected error. name>: <volume-name>
failed fatally.
42059 Invalid set of parameters that Production Run on objects Contact Customer Support.
are passed to discovery did not run. Errors occurred:
export policy. The following parameters
are missing: action_limit.
(42059)
42060 Discovery export policy failed Production Run on objects Verify that the discovery
to create target directory for (Copying native objects) export volume has write
the export. failed, unable to create permission and re-execute
target dir: production/10. policy.
(42060)
42062 Discovery export policy was Production Run on objects Contact Customer Support.
ended abnormally. (Copying native objects)
terminated abnormally.
(42062)
42088 The full-text optimization Full-text optimization failed Contact Customer Support.
process failed; however, the on volume <volume-name>
index is most likely still (42088)
usable for queries.

Event log messages 167


Table 65. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
45802 A full-text index is already Time allocated to gain Contact Customer Support.
being modified. exclusive access to in-
memory index for volume=
1357 has expired (45802)
45803 The index for the specified Index '/deepfs/full-text/ No user intervention is
volume does not exist. This volume_index/ required.
message can occur under volume_1357' not found.
normal conditions. (45803)
45804 Programming error. A Transaction of client: Contact Customer Support.
transaction was never [Link].com_ FINDEX_
initiated or was closed early. QUEUE_ 1357_117251522
_3_ 2 is not the writer
(45804)
45805 The query is not started or Query ID: 123 does not exist No user intervention is
expired. The former is a (45805) required.
programming error. The latter
is normal.
45806 The query expression is Failed to parse 'dog pre\3 Revise your full-text query.
invalid or not supported. bar' (45806)
45807 Programming error. A Client: [Link].com_ Contact Customer Support.
transaction was already FINDEX_QUEUE
started for the client. _1357_1172515 222_3_2
is already active (45807)
45808 A transaction was never No transaction for client: No user intervention is
started or expired. [Link].com_ FINDEX_ required. The system
QUEUE_1357_ handles this condition
1172515222_3_2 (45808) internally.
45810 Programming error. Invalid volumeId. Expected: Contact Customer Support.
1357 Received:2468
(45810)
45812 A File I/O error occurred Failed to write disk (45812). Try your query again.
while the system was Contact Customer Support
accessing index data. for more assistance if
necessary.
45814 The query expression is too Query: 'a* b* c* d* e*' is too Refine your full-text query.
long. complex (45814)
45815 The file that is being indexed Java heap exhausted while Check the skipped file list in
is too large or the query indexing node with ID: the audit log for files that
expression is too complex. '10f4179cd5ff22f 2a6b failed to load due to their
The engine temporarily ran 79a1bc3aef247 fd94ccff' sizes. Revise your query
out of memory. (45815) expression and try again.
46023 Tar command failed while it Failed to back up full-text Check disk space and
persists full-text data to data for server:share. permissions.
Windows Share or NFS share. Reason: <reason>. (46023)

168 IBM StoredIQ: Data Server Administration Guide


Table 65. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
46024 Unhandled unrecoverable Exception <exception> Contact Customer Support.
exception while persisting while backing up fulltext
full-text data into a .tgz file. data for server:share
(46024)
46025 Was not able to delete Failed to unlink incomplete Check permissions.
partial .tgz file after a failed backup image. Reason:
full-text backup. <reason>. (46025)
47002 Synchronization failed on a Synchronization failed for Contact Customer Support.
query. query '<query-name>' on
volume '<server-and-
volume> (47002)
47101 An error occurred during the Cannot process full-text Restart services and contact
query of a full-text expression (Failed to read Customer Support.
expression. from disk (45812) (47101)
47203 No more database Database connections Contact Customer Support.
connections are available. exhausted (512/511)
(47203)
47207 User is running out of disk Disk usage exceeds Contact Customer Support.
space. threshold. (%d) In rare cases, this message
can indicate a program error
leaking disk space. In most
cases, however, disk space
is almost full, and more
storage is required.
47212 Interrogator failed while the Harvester 1 Does not exist. If the problem persists (that
system processed a file. The Action taken : restart. is, the system fails on the
current file is missing from (47212) same file or type of files),
the volume cluster. contact Customer Support.
47214 SNMP notification sender is Unable to resolve host name Check spelling and DNS
unable to resolve the trap [Link] [Link] setup.
host name. (47214)
50011 The DDL/DML files that are Database version control Contact Customer Support.
required for the database SQL file not found. (50011)
versioning were not found in
the expected location on the
data server.
50018 Indicates that the pre- Database restore is Contact Customer Support.
upgrade database restoration unsuccessful. Contact
failed, which was attempted Customer Support. (50018)
as a result of a database
upgrade failure.
50020 Indicates that the current Versions do not match! Contact Customer Support.
database requirements do Expected current database
not meet those requirements version: <dbversion>.
that are specified for the (50020)
upgrade and cannot proceed
with the upgrade.

Event log messages 169


Table 65. ERROR event log messages. This table lists all ERROR event-log messages. (continued)
Event
number Reason Sample message Required customer action
50021 Indicates that the full Database backup failed. Contact Customer Support.
database backup failed when (50021)
the system attempts a data-
object level database backup.
61003 Discovery export policy failed Production policy failed to
to mount volume. mount volume. Aborting.
(61003)
61005 The discovery export load file Production load file Contact Customer Support.
generation fails generation failed. Load files
unexpectedly. The load files may be produced, but post-
can be produced correctly, processing may be
but post-processing actions incomplete. (61005)
like updating audit trails and
generating report files might
not complete.
61006 The discovery export load file Production load file Free up space on the target
generation was interrupted generation interrupted. disk, void the discovery
because the target disk is full. Target disk full. (61006) export run and run the
policy again.
68001 The gateway and data server Gateway connection failed Update your data server to
must be on the same version due to unsupported data the same build number as
to connect. server version. the gateway and restart
services. If your encounter
issues, contact Customer
Support.
68003 The data server failed to The data-server connection Contact Customer Support.
connect to the gateway over to the gateway cannot be
an extended period. established.
8002 The system failed to open a Failed to connect to the The "maximum database
connection to the database. database (80002) connections" configuration
parameter of the database
engine might need to be
increased. Contact
Customer Support.

INFO event log messages


The following table contains a complete listing of all INFO event-log messages.

Table 66. INFO event log messages


Event Required customer
number Reason Sample message action

9001 No conditions were added to Harvester: Query <query name> Add conditions to the
a query. cannot be inferred because no specified query.
condition for it has been defined
(9001).

170 IBM StoredIQ: Data Server Administration Guide


Table 66. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

9002 One or more conditions in a Harvester: Query < query name> Verify that regular
query were incorrect. cannot be inferred because of expressions are
regular expression or other properly formed.
condition error (9002).

9003 Volume Harvest is complete Volume statistics computation No user intervention is


and explorers are being started (9003). required.
calculated.

9004 Explorer calculations are Volume statistics computation No user intervention is


complete. completed (9004). required.

9005 Query membership Query inference will be done in No user intervention is


calculations started. <number> steps (9005). required.

9006 Query membership Query inference step <number> No user intervention is


calculations progress done (9006). required.
information.

9007 Query membership Query inference completed (9007). No user intervention is


calculations completed. required.

9012 Indicates the end of Dump of Volume cache(s) No user intervention is


dumping the content of the completed (9012). required.
volume cache.

9013 Indicates the beginning of Postprocessing for volume No user intervention is


the load process. 'Company Data Server:/mnt/demo- required.
A' started (9013).

9067 Indicates load progress. System metadata and tagged values No user intervention is
were successfully loaded for volume required.
'server:volume' (9067).

9069 Indicates load progress. Volume 'data server: /mnt/demo-A': No user intervention is
System metadata, tagged values required.
and full-text index were
successfully loaded (9069).

9084 The volume export finished. Exporting volume 'data server:/mnt/ No user intervention is
demo-A' (1357) completed (9084) required.

9087 The volume import finished. Importing volume 'dataserver:/mnt/ No user intervention is
demo-A' (1357) completed (9087) required.

9091 The load process was ended Load aborted due to user request No user intervention is
by the user. (9091). required.

15008 The volume load step was Post processing skipped for volume No user intervention is
skipped, per user request. <server>:<vo-lume>. (15008) required.

Event log messages 171


Table 66. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

15009 The volume load step was Harvest skipped for volume No user intervention is
run but the harvest step was <server>:<vol-ume>. (15009) required.
skipped, per user request.

15012 The policy that ran on the Volume <volume> on server No user intervention is
volume is complete and the <server> is free now. Proceeding required.
volume load can now with load. (15012)
proceed.

15013 The configured time limit on Harvest time limit reached for No user intervention is
a harvest was reached. server:share. Ending harvest now. required.
(15013)

15014 The configured object count Object count limit reached for No user intervention is
limit on a harvest was server:share. Ending harvest now. required.
reached. (15014)

15017 checkbox is selected for Deferring post processing for No user intervention is
nightly load job. volume server:vol (15017) required.

15018 Harvest size or time limit is Harvest limit reached on No user intervention is
reached. server:volume. Synthetic deletes required.
will not be computed. (15018)

15019 User stops harvest process. Harvest stopped by user while No user intervention is
processing volume dpfsvr:vol1. Rest required.
of volumes will be skipped. (15019)

15020 The harvest vocabulary Vocabulary for dpfsvr:jhaide-A has Full harvest must be run
changed. Full harvest must changed. A full harvest is instead of an
run instead of incremental. recommended (15020). incremental harvest.

15022 The user is trying to run an Permission-only harvest: permission No action is needed as
ACL-only harvest on a checks not supported for the volume is skipped.
volume that is not a <server>:<share>
Windows Share or
SharePoint volume.

15023 The user is trying to run an Permission-only harvest: volume No action is needed as
ACL-only harvest on a <server>:<share> has no associated the volume is skipped.
volume that was assigned a user list.
user list.

17507 Limit (time or object count) Retention discovery limit reached Contact Customer
reached for generic retention for <server>:<share> Support.
discovery.

172 IBM StoredIQ: Data Server Administration Guide


Table 66. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

17508 Generic retention discovery No new items discovered. Post- No user intervention is
found no new items for this processing skipped for volume required unless the
master volume. <server>:<share> user is certain that new
items must be
discovered.

17509 Generic retention discovery Created new discovered volume No user intervention is
created a new volume. <server>:<share> in volume set required.
<autodis-covered volume set
name>.

18004 Job was stopped. Walker._process File: Grazer No user intervention is


Stopped. (18004) required.

18005 Grazer queue was closed. Walker._processFile: Grazer Closed. No user intervention is
(18005) required.

18016 Displays the list of top-level Choosing top-level directories: No user intervention is
directories that are selected <directories> (18016) required.
by matching the start
directory regular expression.
Displays at the beginning of
a harvest.

34001 Marks current progress of a <volume>: <count> data objects No user intervention is
copy action. processed by copy action. (34001) required.

34004 Marks current progress of a <volume>: <count> data objects No user intervention is
delete action. processed by delete action. (34004) required.

34008 Marks current progress of a <volume>: <count> data objects No user intervention is
move action. processed by move action. (34008) required.

34014 A policy audit was deleted. Deleting Policy Audit # <audit id> No user intervention is
<policy name> <start time> (34014) required.

34015 A policy audit was deleted. Deleted Policy Audit # <audit id> No user intervention is
<policy name> <start time> (34015) required.

34031 Progress update of the Winserver:top share : 30000 data No user intervention is
discovery export policy, objects processed by production required.
every 10000 objects action. (34031)
processed.

41001 A job was started either <jobname> started. (41001) No user intervention is
manually or was scheduled. required.

41002 The user stopped a job that <jobname> stopped at user request No user intervention is
was running. (41002) required.

41003 A job is completed normally <jobname> completed. (41003) No user intervention is


with or without success. required.

Event log messages 173


Table 66. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

41006 Rebooting or restarting Service shutdown. Stopping Rerun jobs after restart
services on the controller or outstanding jobs. (41006) if you want the jobs to
compute node causes all complete.
jobs to stop.

41008 Database compactor Database compactor was not run Set the database
(vacuum) job cannot run because other jobs are active compactor's job
while there is database (41008). schedule so that it does
activity. not conflict with long-
running jobs.

42005 The action completed or was Copy complete: <number> data No user intervention is
ended. Shows results of objects copied,<number> collisions required.
copy action. found. (42005)

42010 The action completed or was Move complete: <number> data No user intervention is
ended. Shows results of objects moved,<number > collisions required.
move action. found.
(42010)

42018 The action completed or was Copy data objects complete: No user intervention is
ended. Shows results of <number> data objects required.
deleted action. copied,<number> collisions found.
(42018)

42024 The synchronizer was Content Data Synchronizer No user intervention is


completed normally. complete. (42024) required.

42028 The action completed or was Policy completed (42028). No user intervention is
ended. Shows results of required.
policy action.

42032 The action completed or was <report name> completed (42032). No user intervention is
ended. Shows results of required.
report action.

42033 The synchronizer started Content Data Synchronizer started. No user intervention is
automatically or manually (42033) required.
with the GUI button.

42048 Reports that the Content Data Synchronizer skipping No user intervention is
synchronizer is skipping a <server-name>:<volume-name> as required.
volume if synchronization is it does not need synchronization.
determined not to be 42048)
required.

42049 Reports that the Content Data Synchronizer starting No user intervention is
synchronizer started synchronization for volume <server- required.
synchronization of a volume. name>:<volume-name>

174 IBM StoredIQ: Data Server Administration Guide


Table 66. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

42053 The policy that was waiting Proceeding with execution of No user intervention is
for participant volumes to be <policy-name>. required.
loaded before it continues, is
now starting.

42063 Report on completion of Production Run on objects (Copying No user intervention is


discovery export policy native objects) completed: 2003 required.
execution phase. data objects copied, 25 duplicates
found. (42063)

42065 A discovery export policy Proceeding with execution of No user intervention is


that was held up for want of 'Production case One'. (42065) required.
resources, is now done
waiting, and begins
execution.

42066 A new discovery export run New run number 10 started for Note the new run
started. production Production Case 23221. number to tie the
(42066) current run with the
corresponding audit
trail.

42067 Discovery export policy is Production Run producing Audit No user intervention is
preparing the audit trail in Trail XML. (42067) required.
XML format. It might take a
few minutes.

42074 A query or tag was replicated Successfully sent query 'Custodian: No user intervention is
to a member data server Joe' to member data server San required.
successfully. Jose Office (42074)

46001 The backup process began. Backup Process Started. (46001) No user intervention is
Any selected backups in the required.
system configuration screen
are run if necessary.

46002 The backup process did not Backup Process Failed: <error- Check your backup
complete all its tasks description>. (46002) volume.
successfully. One or more
backup types did not occur.

46003 The backup process Backup Process Finished. (46003) No user intervention is
completed attempting all the required.
necessary tasks
successfully. Any parts of
the overall process add their
own log entries.

Event log messages 175


Table 66. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

46004 The Application Data Application Data backup failed. Check your backup
backup, as part of the overall (46004) volume. Look at the
backup process, needed to setup for the
run but did not succeed. Application Data
backup. If backups
continue to fail, contact
Customer Support.

46005 The Application Data Application Data backup finished. No user intervention is
backup, as part of the overall (46005) required.
backup process, needed to
run and succeeded.

46006 The Application Data Application Data backup not No user intervention is
backup, as part of the overall configured, skipped. (46006) required.
backup process, was not
configured.

46007 The Harvested Volume Data Harvested Volume Data backup Check your backup
backup, as part of the overall failed. (46007) volume. Look at the
backup process, needed to setup for the Harvested
run but did not succeed. Volume Data backup. If
backups continue to
fail, contact Customer
Support.

46008 The Harvested Volume Data Harvested Volume Data backup No user intervention is
backup, as part of the overall finished. (46008) required.
backup process, needed to
run and succeeded.

46009 The Harvested Volume Data Harvested Volume Data backup not No user intervention is
backup, as part of the overall configured, skipped. (46009) required.
backup process, was not
configured.

46010 The System Configuration System Configuration backup failed. Check your backup
backup, as part of the overall (46010) volume. Look at the
backup process, needed to setup for the System
run but did not succeed. Configuration backup. If
backups continue to
fail, contact Customer
Support.

46011 The System Configuration System Configuration backup No user intervention is


backup, as part of the overall finished. (46011) required.
backup process, needed to
run and succeeded.

176 IBM StoredIQ: Data Server Administration Guide


Table 66. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

46012 The System Configuration System Configuration backup not No user intervention is
backup, as part of the overall configured, skipped. (46012) required.
backup process, was not
configured.

46013 The Audit Trail backup, as Policy Audit Trail backup failed. Check your backup
part of the overall backup (46013) volume. Look at the
process, needed to run but setup for the Audit Trail
did not succeed. backup. If back-ups
continue to fail, contact
IBM support.

46014 The Audit Trail backup, as Policy Audit Trail backup finished. No user intervention is
part of the overall backup (46014) required.
process, needed to run and
succeeded.

46015 The Audit Trail backup, as Policy Audit Trail backup not No user intervention is
part of the overall backup configured, skipped. (46015) required.
process was not configured.

46019 Volume cluster backup Indexed Data backup failed: Contact Customer
failed. <specific error> (46019) Support.

46020 Volume cluster backup Indexed Data backup finished. No user intervention is
finished. (46020) required.

46021 Volume is not configured for Indexed Data backup not No user intervention is
indexed data backups. configured, skipped. (46021) required.

46022 Full-text data was Successfully backed up full-text No user intervention is


successfully backed up. data for server:share (46022) required.

47213 Interrogator was Harvester 1 is now running. (47213) No user intervention is


successfully restarted. required.

60001 The user updates an object Query cities was updated by the No user intervention is
on the system. It includes administrator account (60001). required.
any object type on the data
server, including the
updating of volumes.

60002 The user creates an object. Query cities was created by the No user intervention is
It includes any object type administrator account (60002). required.
on the data server, including
the creation of volumes.

60003 The user deletes an object. Query cities was deleted by the No user intervention is
It includes any object type administrator account (60003). required.
on the data server, including
the deletion of volumes.

Event log messages 177


Table 66. INFO event log messages (continued)
Event Required customer
number Reason Sample message action

60004 The user publishes a full-text Query cities draft was published by No user intervention is
query set or a query. the administrator account (60004). required.

60005 The user tags an object. It Query tagging for cities class was No user intervention is
includes a published query, a started by the administrator account required.
draft query, or tag. (60005).

60006 A user restarted services on Application services restart for all No user intervention is
the data server. data servers was requested by the required.
administrator account (60006).

61001 Concordance discovery Preparing for upload of load file(s). No user intervention is
export is now preparing the (61001) required.
load files.

61002 Concordance discovery Load file(s) ready for upload. No user intervention is
export is ready to upload the (61002) required.
load files.

65000 The log file finished Log file download complete (65000) No user intervention is
downloading. required.

WARN event log messages


The following table contains a complete listing of all WARN event-log messages, reasons for occurrence,
sample messages, and any required customer action.

Table 67. WARN event log messages


Event
number Reason Sample message Customer action

1002 An Interrogator process failed Processing could not be Classify the document
because of an unknown error. completed on object, manually and Contact
The data object that was interrogator died : <data Customer Support.
processing is skipped. A new object name>. (1002)
process is created to replace
it.

1003 Interrogator child process did Interrogator terminated Try to readd the volume that
not properly get started. before accessing data is harvested. If that fails,
There might be problems to objects. (1003) contact Customer Support.
access the volume to be
harvested.

1004 Interrogator child process Processing was not Contact Customer Support.
was ended because it was no completed on object,
longer responding. The data interrogator killed : <data
object that was processing is object name>. (1004)
skipped. A new process is
created to replace it.

178 IBM StoredIQ: Data Server Administration Guide


Table 67. WARN event log messages (continued)
Event
number Reason Sample message Customer action

6001 A user email might not be Failed to send an email to Verify that your SMTP server
sent. The mail server settings user <email address>; check is configured correctly. Make
are incorrect. mail server configuration sure that the IP address that
settings (6001). is configured for the data
server can relay on the
configured SMTP server.

8001 The database needs to be The Database is approaching Run the database
vacuumed. an operational limit. Please maintenance task to vacuum
run the Database the database.
maintenance task using the
Console interface (8001)

9068 Tagged values were loaded, System metadata and tagged Contact Customer Support.
but full-text index loading values were loaded
failed. successfully for volume
'server:volume', but loading
the full-text index failed
(9068)

9070 Tagged values and full-text Loading system metadata, Contact Customer Support.
index loading failed. tagged values and the full-
text index failed for volume
'server:volume' (9070)

15003 The volume mount appeared Volume <volume name> on Contact Customer Support.
to succeed, but the test for server <server name> is not
mount failed. mounted. Skipping. (15003)

15004 A component cleanup failure [<component>] Cleanup Contact Customer Support.


on stop or completion. failure on stop. (15004)

15005 There was a component run [<component>] Run failure. Contact Customer Support.
failure. (15005)

15006 Cleanup failed for component [<component>] Cleanup Contact Customer Support.
after a run failure. failure on abort. (15006)

15007 A component that is timed Component [<component>] Try your action again. If this
out needs to be stopped. unresponsive; autostopping error continues, contact
triggered. (15007) Customer Support.

15010 The same volume cannot be Volume <volume-name> on No user intervention is


harvested in parallel. The server <server-name> is required. You might want to
harvest is skipped and the already being harvested. verify that the volume harvest
next one, if any are in queue, Skipping. (15010) is complete.
started.

Event log messages 179


Table 67. WARN event log messages (continued)
Event
number Reason Sample message Customer action

15011 A volume cannot be Volume <volume-name> on No user intervention is


harvested if it is being used server <server-name> is required.
by another job. The harvest being used by another job.
continues when the job is Waiting before proceeding
complete. with load. (15011)

15015 Configured harvest time limit Time limit for harvest Reconfigure harvest time
is reached. reached. Skipping Volume v1 limit.
on server s1. 1 (15015)

15016 Configured harvest object Object count limit for harvest Reconfigure harvest data
count limit is reached. reached. Skipping Volume v1 object limit.
on server s1 (15016)

17008 Query that ran to discover Centera External Iterator : Contact Customer Support.
Centera items ended Centera Query terminated
unexpectedly. unexpectedly (<error
description>). (17008)

17011 Running discovery on the Pool Jpool appears to have Make sure that two jobs are
same pool in parallel is not another discovery running. not running at the same time
allowed. Skipping. (17011). that discovers the same pool.

17502 Generic retention discovery is Volume <server>:<share> No user intervention is


already running for this appears to have another required as the next step, if
master volume. discovery running. Skipping. any, within the job is run.

17504 Sent when a retention Volume <server>:<share> is Contact Customer Support.


discovery is run on any not supported for discovery.
volume other than a Windows Skipping.
Share retention volume.

18007 Directory listing or processing Walker._walktree: OSError - Make sure that the appliance
of data object failed in Grazer. <path><rea-son> (18007) still has appropriate
permissions to a volume.
Verify that there is network
connectivity between the
appliance and your volume.
Contact Customer Support.

18008 Unknown error occurred Walker._walktree: Unknown Contact Customer Support.


while processing data object exception - <path>. (18008)
or listing directory.

18009 Grazer timed out processing Walker._process File: Grazer Contact Customer Support.
an object. Timed Out. (18009)

18010 The skipdirs file is either not Unable to open skipdirs file: Contact Customer Support.
present or not readable by <filename>. Cannot skip
root. directories as configured.
(18010)

180 IBM StoredIQ: Data Server Administration Guide


Table 67. WARN event log messages (continued)
Event
number Reason Sample message Customer action

18011 An error occurred reading the Grazer._run: couldn't read Contact Customer Support.
known extensions list from extensions - <reason>.
the database. (18011)

18012 An unknown error occurred Grazer._run: couldn't read Contact Customer Support.
reading the known extensions extensions. (18012)
list from the database.

18015 NFS initialization warning that NIS Mapping not available. User name and group names
NIS is not available. (18015) might be inaccurate. Check
that your NIS server is
available and properly
configured in the data server.

18019 The checkpoint that is saved Unable to load checkpoint for If the message repeats in
from the last harvest of the NewsGator volume. A full subsequent harvests, contact
NewsGator data source failed harvest will be performed Customer Support.
to load. Instead of conducting instead. (18019)
an incremental harvest, a full
harvest is run.

18020 The checkpoint noted for the Unable to save checkpoint for If the message repeats in
current harvest of the NewsGator harvest of subsequent harvests, contact
NewsGator data source might volume. (18020) Customer Support.
not be saved. The next
incremental harvest of the
data source is not able to pick
up from this checkpoint.

33016 System might not unmount Windows Share Protocol Server administrators can see
this volume. Session teardown failed. that connections are left
(33016) hanging for a predefined time.
These connections will drop
off after they time out. No
user intervention required.

33017 System encountered an error An error occurred while Contact Customer Support.
while it tries to figure out retrieving the query instances
what query uses this volume. pointing to a volume. (33017)

33028 The tear-down operation of IBM FileNet tear-down None


the connection to a FileNet operation failed. (33028)
volume failed. Some
connections can be left open
on the FileNet server until
they time out.

34003 Skipped a copy data object Copy action error :- Target Verify that there is space
because disk full error. disk full, skipping copy : available on your policy
<source volume> to <target destination and try again.
volume>. (34003)

Event log messages 181


Table 67. WARN event log messages (continued)
Event
number Reason Sample message Customer action

34010 Skipped a move data object Move action error :- Target Verify that there is space
because disk full error. disk full, skipping copy : available on your policy
<source volume> to <target destination. After verifying
volume>. (34010) that space is available, run
another harvest before you
run your policy. Upon harvest
completion, try running the
policy again.

34029 Discovery export policy Discovery export Run action Create sufficient space on
detects the target disk is full error: Target disk full, target disk and run discovery
and skips production of an skipping discovery export: export policy again.
object. share-1/saved/
[Link] to
production/10/
documents/1/0x0866e
5d6c898d9ffdbea 720b0
90a6f46d3058605 .txt.
(34029)

34032 The policy that was run has No volumes in scope for Check policy query and
no volumes in scope that is policy. Skipping policy scoping configuration, and re-
based on the configured execution. (34032) execute policy.
query and scoping. The policy
cannot be run.

34035 If the global hash setting for Copy objects : Target hash If target hashes need to be
the system is set to not will not be computed because computed for the policy audit
compute data object hash, no Hashing is disabled for trail, turn on the global hash
hash can be computed for the system. (34035) setting before you run the
target objects during a policy policy.
action.

34036 The policy has no source The policy has no source Confirm that the query used
volumes in scope, which volume(s) in scope. Wait for by the policy has one or more
means that the policy cannot the query to update before volumes in scope.
be run. executing the policy. (34036)

42003 The job in this action is Copy data objects stopped at No user intervention is
stopped by the user. user request. (42003) required.

42008 The job in this action is Move data objects stopped at No user intervention is
stopped by the user. user request. (42008) required.

42016 The job in this action is Delete data objects stopped No user intervention is
stopped by the user. at user request. (42016) required.

42026 The job in this action is Policy stopped at user No user intervention is
stopped by the user. request. (42026) required.

182 IBM StoredIQ: Data Server Administration Guide


Table 67. WARN event log messages (continued)
Event
number Reason Sample message Customer action

42035 When the job in this action is Set security for data objects No user intervention is
stopped by the user. stopped at user request. required.
(42035)

42051 Two instances of the same Policy <policy-name> is No user intervention is


policy cannot run at the same already running. Skipping. required.
time. (42051)

42052 Policies cannot be run until One or more volume(s) No user intervention is
after volumes are loaded, if needed by policy <policy- required.
those volumes are name> are being loaded.
participants to the policy by Waiting for the bulk load(s) to
virtue of being in the query. finish. (42052)

42061 Discovery export policy was Discovery export run on No user intervention is
stopped by user. objects (Copying native required.
objects) stopped at user
request. (42061)

42064 Discovery export policy A Discovery export run The discovery export policy
execution is delayed because related to policy 'Discovery execution is held up for
a conflicting discovery export export case One' is in required resources. Execution
run is in progress. progress. Waiting for it to must begin as soon as
finish. (42064) resource becomes available.

42068 Policy failed to set Copy objects warning, unable Policy might not be able to
appropriate permissions on to set permissions on target set appropriate permissions
the target directory. Objects directory: share-1/saved. on the objects it creates. If it
that are created from the (42068) is not acceptable, verify that
policy might not have target volume has
appropriate permissions set. appropriate write
permissions and re-execute.

42069 If the “Copy data objects Discovery export DAT_Export If the modified objects need
modified since last harvest” is configured to act on to be acted upon, either use a
option is selected for a members of containers, and discovery export action only
discovery export policy, it is cannot act on objects on the original file/email
valid only if the discovery modified after the last archive, or conduct an
export itself is defined to act harvest. Discovery export run incremental harvest on the
on the original file/email X will skip modified objects. source volumes.
archive, as opposed to their (42069)
members. If it is not true, the
warning tells the user that
modified objects are skipped.

46026 Volume is being harvested or Volume volume:share is in Rerun backup when volume is
policies are running against it. use. Unable to back up full- not in use.
If there are other full-text text index. Will retry later.
indexes to be backed up, the (46026)
system works on those
actions. Try this volume
again.

Event log messages 183


Table 67. WARN event log messages (continued)
Event
number Reason Sample message Customer action

47201 Database connections are Database connections at No user intervention is


down to a normal level. normal level again (512/100) required.
(47201)

47202 The system is starting to run Database connections usage Contact Customer Support.
low on database connections. seems excessive (512/415)
This situation is abnormal. An (47202)
indication of process restarts
and connections are not
being cleared.

47215 Someone internally or SSHD: Failed password for Contact your local IT
externally is trying (and root from [Link] port manager. It might be either a
failing) to SSH into the data 57982. (47125) mistyped password by a
server. legitimate user or in the worst
case scenario, a genuine
break-in attempt.

61003 One of the load files cannot Failed to mount transaction Some of the load files will be
be uploaded because the cache dump '/deepfs/ missing after the discovery
compute node might not be postgres/pro-duction_cache'. export completes. These load
accessed to obtain. (61003) files are reproduced on a new
run. If problem persists
across runs, Contact
Customer Support.

61004 Warns the user that one of Transaction Cache Dump Run the discovery export
the transaction cache failed with error - Validation policy that saw the error
memory dump processes failed during creation of load again. If the error persists,
encountered an error. If a file. (61004) and you cannot find any
discovery export runs, it cluster/data server
means that the discovery configuration issues, contact
export fails to produce one of Customer Support.
the load files.
Note: If multiple memory
dumps fail, there is one
warning per failed memory
dump.

184 IBM StoredIQ: Data Server Administration Guide


Notices
This information was developed for products and services offered in the U.S.A. This material may be
available from IBM in other languages. However, you may be required to own a copy of the product or
product version in that language in order to access it.
IBM may not offer the products, services, or features discussed in this document in other countries.
Consult your local IBM representative for information on the products and services currently available in
your area. Any reference to an IBM product, program, or service is not intended to state or imply that only
that IBM product, program, or service may be used. Any functionally equivalent product, program, or
service that does not infringe any IBM intellectual property right may be used instead. However, it is the
user's responsibility to evaluate and verify the operation of any non-IBM product, program, or service.
IBM may have patents or pending patent applications covering subject matter described in this
document. The furnishing of this document does not grant you any license to these patents. You can send
license inquiries, in writing, to:

IBM Director of Licensing


IBM Corporation
North Castle Drive
Armonk, NY 10504-1785
U.S.A.

For license inquiries regarding double-byte (DBCS) information, contact the IBM Intellectual Property
Department in your country or send inquiries, in writing, to:

Intellectual Property Licensing


Legal and Intellectual Property Law
IBM Japan Ltd.
19-21, Nihonbashi-Hakozakicho, Chuo-ku
Tokyo 103-8510, Japan

INTERNATIONAL BUSINESS MACHINES CORPORATION PROVIDES THIS PUBLICATION "AS IS"


WITHOUT WARRANTY OF ANY KIND, EITHER EXPRESS OR IMPLIED, INCLUDING, BUT NOT LIMITED TO,
THE IMPLIED WARRANTIES OF NON-INFRINGEMENT, MERCHANTABILITY OR FITNESS FOR A
PARTICULAR PURPOSE. Some jurisdictions do not allow disclaimer of express or implied warranties in
certain transactions, therefore, this statement may not apply to you.
This information could include technical inaccuracies or typographical errors. Changes are periodically
made to the information herein; these changes will be incorporated in new editions of the publication.
IBM may make improvements and/or changes in the product(s) and/or the program(s) described in this
publication at any time without notice.
Any references in this information to non-IBM Web sites are provided for convenience only and do not in
any manner serve as an endorsement of those Web sites. The materials at those Web sites are not part of
the materials for this IBM product and use of those Web sites is at your own risk.
IBM may use or distribute any of the information you supply in any way it believes appropriate without
incurring any obligation to you.
Licensees of this program who wish to have information about it for the purpose of enabling: (i) the
exchange of information between independently created programs and other programs (including this
one) and (ii) the mutual use of the information which has been exchanged, should contact:

IBM Director of Licensing


IBM Corporation
North Castle Drive, MD-NC119
Armonk, NY 10504-1785
US

© Copyright IBM Corp. 2001, 2020 185


Such information may be available, subject to appropriate terms and conditions, including in some cases,
payment of a fee.
The licensed program described in this document and all licensed material available for it are provided by
IBM under terms of the IBM Customer Agreement, IBM International Program License Agreement or any
equivalent agreement between us.
The performance data discussed herein is presented as derived under specific operating conditions.
Actual results may vary.
Information concerning non-IBM products was obtained from the suppliers of those products, their
published announcements or other publicly available sources. IBM has not tested those products and
cannot confirm the accuracy of performance, compatibility or any other claims related to non-IBM
products. Questions on the capabilities of non-IBM products should be addressed to the suppliers of
those products.
Statements regarding IBM's future direction or intent are subject to change or withdrawal without notice,
and represent goals and objectives only.
This information contains examples of data and reports used in daily business operations. To illustrate
them as completely as possible, the examples include the names of individuals, companies, brands, and
products. All of these names are fictitious and any similarity to the names and addresses used by an
actual business enterprise is entirely coincidental.
COPYRIGHT LICENSE:
This information contains sample application programs in source language, which illustrate programming
techniques on various operating platforms. You may copy, modify, and distribute these sample programs
in any form without payment to IBM, for the purposes of developing, using, marketing or distributing
application programs conforming to the application programming interface for the operating platform for
which the sample programs are written. These examples have not been thoroughly tested under all
conditions. IBM, therefore, cannot guarantee or imply reliability, serviceability, or function of these
programs. The sample programs are provided "AS IS", without warranty of any kind. IBM shall not be
liable for any damages arising out of your use of the sample programs.
Each copy or any portion of these sample programs or any derivative work, must include a copyright
notice as follows:
© (your company name) (year).
Portions of this code are derived from IBM Corp. Sample Programs.
© Copyright IBM Corp. _enter the year or years_.

Trademarks
IBM, the IBM logo, and [Link] are trademarks or registered trademarks of International Business
Machines Corp., registered in many jurisdictions worldwide. Other product and service names might be
trademarks of IBM or other companies. A current list of IBM trademarks is available on the Web at
"Copyright and trademark information" [Link]
Adobe and PostScript are either registered trademarks or trademarks of Adobe Systems Incorporated in
the United States, and/or other countries.
Microsoft and Windows are trademarks of Microsoft Corporation in the United States, other countries, or
both.
Java and all Java-based trademarks and logos are trademarks or registered trademarks of Oracle and/or
its affiliates.
UNIX is a registered trademark of The Open Group in the United States and other countries.
VMware, VMware vCenter Server, and VMware vSphere are registered trademarks or trademarks of
VMware, Inc. or its subsidiaries in the United States and/or other jurisdictions.

186 Notices
The registered trademark Linux is used pursuant to a sublicense from the Linux Foundation, the exclusive
licensee of Linus Torvalds, owner of the mark on a worldwide basis.
Red Hat and OpenShift are trademarks or registered trademarks of Red Hat, Inc. or its subsidiaries in the
United States and other countries.

Terms and conditions for product documentation


Permissions for the use of these publications are granted subject to the following terms and conditions.

Applicability
These terms and conditions are in addition to any terms of use for the IBM website.

Personal use
You may reproduce these publications for your personal, noncommercial use provided that all proprietary
notices are preserved. You may not distribute, display or make derivative work of these publications, or
any portion thereof, without the express consent of IBM.

Commercial use
You may reproduce, distribute and display these publications solely within your enterprise provided that
all proprietary notices are preserved. You may not make derivative works of these publications, or
reproduce, distribute or display these publications or any portion thereof outside your enterprise, without
the express consent of IBM.

Rights
Except as expressly granted in this permission, no other permissions, licenses or rights are granted, either
express or implied, to the publications or any information, data, software or other intellectual property
contained therein.
IBM reserves the right to withdraw the permissions granted herein whenever, in its discretion, the use of
the publications is detrimental to its interest or, as determined by IBM, the above instructions are not
being properly followed.
You may not download, export or re-export this information except in full compliance with all applicable
laws and regulations, including all United States export laws and regulations.
IBM MAKES NO GUARANTEE ABOUT THE CONTENT OF THESE PUBLICATIONS. THE PUBLICATIONS ARE
PROVIDED "AS-IS" AND WITHOUT WARRANTY OF ANY KIND, EITHER EXPRESSED OR IMPLIED,
INCLUDING BUT NOT LIMITED TO IMPLIED WARRANTIES OF MERCHANTABILITY, NON-
INFRINGEMENT, AND FITNESS FOR A PARTICULAR PURPOSE.

IBM Online Privacy Statement


IBM Software products, including software as a service solutions, (“Software Offerings”) may use cookies
or other technologies to collect product usage information, to help improve the end user experience, to
tailor interactions with the end user or for other purposes. In many cases no personally identifiable
information is collected by the Software Offerings. Some of our Software Offerings can help enable you to
collect personally identifiable information. If this Software Offering uses cookies to collect personally
identifiable information, specific information about this offering’s use of cookies is set forth below.
This Software Offering does not use cookies or other technologies to collect personally identifiable
information.
If the configurations deployed for this Software Offering provide you as customer the ability to collect
personally identifiable information from end users via cookies and other technologies, you should seek

Notices 187
your own legal advice about any laws applicable to such data collection, including any requirements for
notice and consent.
For more information about the use of various technologies, including cookies, for these purposes, See
IBM’s Privacy Policy at [Link] and IBM’s Online Privacy Statement at http://
[Link]/privacy/details the section entitled “Cookies, Web Beacons and Other Technologies” and
the “IBM Software Products and Software-as-a-Service Privacy Statement” at [Link]
software/info/product-privacy.

188 IBM StoredIQ: Data Server Administration Guide


Index

A data object types 22


data object typing 23
administration knobs data server
configuring 134 LDAP connections 16
alias 25 data sources
attributes adding 25
SharePoint 151 delete
audit action 98
saving results from 130 policy 98
audit messages desktop 97
policy 131 Desktop collection 94
audit results desktop services
saving 130 enable 96
audit settings desktop settings
configuring 23 configuring 96
Audit tab desktop volume, deleting 77
Configuration subtab 2 discovery export volume
Dashboard subtab 2 creating 75
Data Sources subtab 2 Documentum
audit trails as a data source 67
searching 129 configuring as server platform 29
audits installing 29
harvest 99 Domino
policy 127 adding as primary volume 65

B E
back up 15 ERROR event log messages 104, 162
backup 14 event
subscribing to 103
event log
C clearing current 103
Chatter downloading 103
configuring messages 66 viewing 103
CIFS event log messages
configuring server platform 26 ERROR 104, 162
CIFS (Windows platforms) 70 INFO 104, 162
CIFS volumes WARN 104, 162
Copy action 64 event logs 103
Move action 64 Exchange
CIFS/SMB or SMB2 30, 154 servers 26
CMIS 30, 154 Exchange 2003
configuring SMB properties 62 improving performance 27
Connections Exchange 2007 Client Access Server support 64
setting up the administrator access on Connections 68 Exchange Admin Role 26
content based hash 23 Exchange servers
content typing 23 configuring 26
Copy action
CIFS volumes F
ownership 64
customized web services 133 file types
supported 135, 143
FileNet
D configuring 66
DA Gateway settings full hash 23
configuring 9 full-text index settings
configure 20

Index 189
H job (continued)
starting 90
harvest jobs
full 81 Centera deleted files synchronizer 89
incremental 81 CIFS/NFS retention volume deleted files synchronizer
lightweight 82 89
monitoring 84 Database Compactor 89
post-processing 81 Harvest every volume 89
troubleshooting 92 System maintenance and cleanup 89
harvest audits Update Age Explorers 89
viewing 101
harvest list
downloading 102
L
harvest tracker 84 LDAP connections
harvester data server 16
lightweight harvest legal
harvester settings 83 notices 185
harvester settings 19 lightweight harvest
harvesting full-text settings 83
about 81 hash settings 83
harvests performing 82
incremental 81 volume configuration 82
pausing 84 Livelink 30
resuming 84 logging in
hash IBM StoredIQ Data Server 2
content based 23 logging out
full 23 IBM StoredIQ Data Server 2
metadata based 23 logs 99
partial 23
settings 23
hash computation M
SharePoint 151
mail settings
HDFS 30
configuring 11
messages
I event log 104, 162
metadata based hash 23
IBM Connections 67 Microsoft Exchange 154
IBM Connections data source 67 migrations 80
IBM Content Manager 30, 69, 154 monitoring
IBM StoredIQ Desktop Data Collector harvest 84
installation 94 Move action
installer 95 CIFS volumes
prerequisites 94 ownership 64
IBM StoredIQ Desktop Data Collector files 98
IBM StoredIQ image 15
IIS 6.0 N
improving performance 27
network settings 10
import audits 102
NewsGator
incremental harvests 81
required privileges 29
INFO event log messages 104, 162
NFS 26, 154
installation methods 95
NFS v2, v3 30
installer
NFS v3 70
downloading from the application 95
notices
interface
legal 185
buttons 4
NSF files
icons 4
importing from Lotus Notes 18

J
O
Jive 30, 154
OCR processing 20
job
Optical Character Recognition 20
creating 89
ownership
predefined 90
CIFS volumes 64

190 IBM StoredIQ: Data Server Administration Guide


ownership (continued) SiqDocument 69
Copy action 64 Skipped directory 30
Move action 64 SMB 62
preserving 64 SMB properties 62
SMB1 62
SMB2 62
P SMBv2 62
partial hash 23 SNMP settings
pausing harvests 84 configuring 11
policy audit supported file types
by name 127 archive 135, 143
by volume 127 CAD 135, 143
viewing details 128 database 135, 143
policy audit failure messages 131 email 135, 143
policy audit messages 131 graphic 135, 143
policy audit success messages 131 multimedia 135, 143
policy audit warning messages 131 presentation 135, 143
policy audits 127 spreadsheet 135, 143
predefined job system 135, 143
running 90 text and markup 135, 143
prerequisites word processing 135, 143
IBM StoredIQ Desktop Data Collector 94 system
primary volume restart 7
Domino 65 system backup configurations 14
primary volumes system restore configurations 14
creating 30 system status
Exchange 2007 Client Access Server support 64 checking 6
processing system volume
monitoring 91 adding 72
protocols
supported 154 T
text extraction 23
R
restore 14 U
resuming harvests 84
retention volume user account
adding 70 administering 17
retention volumes 70 creating 17
deleting 17
locking 17
S resetting passwords 17
search depth unlocking 17
volume indexing 25 user interface
server alias 25 Administration tab 2
server platforms Audit tab 2
supported 154 Folders tab 2
services
restarting 7 V
SharePoint
alternate-access mappings 28 volume
privileges for social data 28 adding primary 30
servers 28 discovery export 75
SharePoint 2013 web services 133 system 72
SharePoint objects volume configuration
supported types 151 desktop 97
SharePoint targets 80 volume data
SharePoint volume exporting 73
performance considerations with versioning 65 exporting to a system volume 73
SharePoint volumes importing 73
performance considerations 65 importing to a system volume 74
permissions for mounting 82 volume definitions
versioning 65 editing 7

Index 191
volume indexing 25
volume-import audits 102
volumes
action limitations by type 77
creating 25
deleting 76
export 154
primary 154
retention 70, 154
system 154

W
WARN event log messages 104, 162
web services
deploy 133
Windows authentication
enabling integration on Exchange servers 27
Windows Share
server 26

192 IBM StoredIQ: Data Server Administration Guide


IBM®

You might also like