0% found this document useful (0 votes)
18 views52 pages

Managing Historical Data Trends

Uploaded by

Surya Basnet
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views52 pages

Managing Historical Data Trends

Uploaded by

Surya Basnet
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Data & Knowledge

Management

Unit 3

12/17/2024 1
Managing Data

• All IT applications require data.


• These data should be of high quality, meaning that they should be:
 accurate,
 complete,
 timely, consistent,
 accessible,
 relevant, and
 Concise (containing lot of clear information)

Unfortunately, the process of acquiring, keeping, and managing data is


becoming
12/17/2024 increasingly difficult. 2
The Difficulties of Managing Data
• Because data are processed in several stages and often in
multiple locations, they are frequently subject to problems
and difficulties.
• Managing data in organizations is difficult for many reasons.
• The amount of data increases exponentially with time.
• Much historical data must be kept for a long time, and new
data are added rapidly.
12/17/2024 3
• For example, to support millions of customers, large retailers
such as Walmart have to manage petabytes of data.
• (A petabyte is approximately 1,000 terabytes, or trillions of
bytes; )
• In addition, data are also scattered throughout organizations,
and they are collected by many individuals using various
methods and devices.
• These data are frequently stored in numerous servers and
locations and in different computing systems, databases,
formats, and human and computer languages.
12/17/2024 4
Data Governance
• To address the numerous problems associated with managing data,
organizations are turning to data governance.
• Data governance is an approach to managing information across an
entire organization.
• It involves a formal set of business processes and policies that are
designed to ensure that data are handled in a certain, well-defined
fashion.
• That is, the organization follows unambiguous rules for creating,
collecting, handling, and protecting its information.
• The objective is to make information available, transparent, and useful for
the people who are authorized to access it, from the moment it enters an
12/17/2024 5
organization until it is outdated and deleted.
• One strategy for implementing data governance is master data
management.
• Master data management is a process that spans all organizational
business processes and applications.
• It provides companies with the ability to store, maintain, exchange, and
synchronize a consistent, accurate, and timely “single version of the
truth” for the company’s master data.
• Master data are a set of core data, such as customer, product,
employee, vendor, geographic location, and so on, that span the
enterprise information systems.
12/17/2024 6
• It is important to distinguish between master data and
transaction data.
• Transaction data, which are generated and captured by
operational systems, describe the business’s activities, or
transactions.
• In contrast, master data are applied to multiple
transactions and are used to categorize, aggregate, and
evaluate the transaction data.

12/17/2024 7
Big Data
We are accumulating data and information at an increasingly rapid pace
from such diverse sources : New sources of data and information
• as company documents, include:
• e-mails, • blogs,
• Web pages, • Social media
• credit card swipes, • podcasts,
• phone messages, stock trades, • videocasts (think of YouTube),
• memos, • digital video surveillance, and
• address books, and • RFID tags and other wireless
• radiology
12/17/2024
scans. sensors
8
• In fact, organizations are capturing data about almost all events—
including events that, in the past, firms never used to think of as
data at all, such as a person’s location, the vibrations and
temperature of an engine, or the stress at numerous points on a
bridge—and then analyzing those data.
• Organizations and individuals must process an unimaginably vast
amount of data that is growing ever more rapidly.
• According to IDC (a technology research fi rm), the world generates
exabytes of data each year (an exabyte is one trillion terabytes).
• Furthermore, the amount of data produced worldwide is increasing
by 50 percent each year.
12/17/2024 9
At its core, Big Data is about predictions. Predictions do not come from
“teaching” computers to “think” like humans. Instead, predictions come
from applying mathematics to huge quantities of data to infer
probabilities.

Consider these examples:


• The likelihood that an e-mail message is spam;
• The likelihood that the typed letters “teh” are supposed to be “the”;
• The likelihood that the trajectory and velocity of a person
jaywalking(unlawful walking) indicates that he will make it across the
street in time—meaning that a self-driving car need only slow down
slightly.
12/17/2024 10
• Big Data systems perform well because they contain
huge amounts of data on which to base their predictions.
• Moreover, they are configured to improve themselves
over time by searching for the most valuable signals and
patterns as more data are input.

12/17/2024 11
Defining Big Data
• It is difficult to defi ne Big Data.
• Here we present two descriptions of the phenomenon.

• First, the technology research firm Gartner


([Link]) defines Big Data as diverse, high-
volume, high-velocity information assets that require
new forms of processing to enable enhanced decision
making, insight discovery, and process optimization.
12/17/2024 12
Second, the Big Data Institute (TBDI;
[Link]) defines Big Data as vast data
sets that:
• Exhibit variety;
• Include structured, unstructured, and semi-structured data;
• Are generated at high velocity with an uncertain pattern;
• Do not fit neatly into traditional, structured, relational
databases and Can be captured, processed, transformed,
and analyzed in a reasonable amount of time only by
sophisticated
12/17/2024
information systems. 13
Big Data generally consists of the following. Keep in mind that this list is
not inclusive. It will expand as new sources of data emerge.
• Traditional enterprise data—examples are customer information
from customer relationship management systems, transactional
enterprise resource planning data, Web store transactions, operations
data, and general ledger data.
• Machine-generated/sensor data—examples are smart meters;
manufacturing sensors; sensors integrated into smartphones,
automobiles, airplane engines, and industrial machines; equipment
logs; and trading systems data.

12/17/2024 14
• Social data—examples are customer feedback
comments; microblogging sites such as Twitter; and
social media sites such as Facebook, YouTube, and
LinkedIn.
• Images captured by billions of devices located
throughout the world, from digital cameras and
camera phones to medical scanners and security
cameras.

12/17/2024 15
Let’s take a look at a few specific examples of Big Data:
• When the Sloan Digital Sky Survey in New Mexico was launched in
2000, its telescope collected more data in its first few weeks than had
been amassed in the entire history of astronomy. By 2013, the survey’s
archive contained hundreds of terabytes of data. However, the Large
Synoptic Survey Telescope in Chile, due to come online in 2016, will
collect that quantity of data every five days.
• In 2013 Google was processing more than 24 petabytes of data every
day.
• Facebook members upload more than 10 million new photos every
hour. In addition, they click a “like” button or leave a comment nearly 3
billion times every day.
12/17/2024 16
• The 800 million monthly users of Google’s YouTube service upload
more than an hour of video every second.
• The number of messages on Twitter grows at 200 percent every
year. By mid-2013 the volume exceeded 450 million tweets per day.
• As recently as the year 2000, only 25 percent of the stored
information in the world was digital. The other 75 percent was
analog; that is, it was stored on paper, film, vinyl records, and the
like. By 2013, the amount of stored information in the world was
estimated to be around 1,200 exabytes, of which less than 2 percent
was non-digital.
12/17/2024 17
Characteristics of Big Data
Big Data has three distinct characteristics: volume, velocity, and
variety. These characteristics distinguish Big Data from traditional
data.
Volume:
• Although the sheer volume of Big Data presents data
management problems, this volume also makes Big Data
incredibly valuable.
• Irrespective of their source, structure, format, and frequency, data
are always valuable.
• If12/17/2024
certain types of data appear to have no value today, 18it is
Velocity:
• The rate at which data flow into an organization is rapidly
increasing.
• Velocity is critical because it increases the speed of the feedback
loop between a company and its customers.
• For example, the Internet and mobile technology enable online
retailers to compile histories not only on final sales, but on their
customers’ every click and interaction.
• Companies that can quickly utilize that information—for example,
by recommending additional purchases—gain competitive
advantage.
12/17/2024 19
Variety:
• Traditional data formats tend to be structured, relatively well
described, and they change slowly.
• Traditional data include financial market data, point-of-sale
transactions, and much more.
• In contrast, Big Data formats change rapidly.
• They include satellite imagery, broadcast audio streams,
digital music files, Web page content, scans of government
documents, and comments posted on social networks.

12/17/2024 20
Managing Big Data
• Big Data makes it possible to do many things that were previously
impossible; for example, spot business trends more rapidly and
accurately, prevent disease, track crime, and so on.
• When properly analyzed, Big Data can reveal valuable patterns and
information that were previously hidden because of the amount of
work required to discover them.
• Leading corporations, such as Walmart and Google, have been able to
process Big Data for years, but only at great expense.

12/17/2024 21
• Today’s hardware, cloud computing , and open- source software
make processing Big Data affordable for most organizations.
• The first step for many organizations toward managing Big Data
was to integrate information silos into a database environment and
then to develop data warehouses for decision making.
• After completing this step, many organizations turned their
attention to the business of information management—making
sense of their proliferating data.
• In recent years, Oracle, IBM, Microsoft, and SAP have spent billions
of dollars purchasing software fi rms that specialize in data
management and business intelligence.
12/17/2024 22
The Database Approach
• From the time that businesses first adopted computer applications
(mid-1950s) until the early 1970s, organizations managed their data
in a fi le management environment.
• This environment evolved because organizations typically automated
their functions one application at a time.
• Therefore, the various automated systems developed independently
from one another, without any overall planning. Each application
required its own data, which were organized in a data file.

12/17/2024 23
• A data file is a collection of logically related records. In a fi le
management environment, each application has a specific data fi le
related to it.
• This file contains all of the data records the application requires.
Over time, organizations developed numerous applications, each
with an associated, application-specific data file.
• Using databases eliminates many problems that arose from previous
methods of storing and accessing data, such as fi le management
systems.

12/17/2024 24
Databases are arranged so that one set of software
programs—the database management system—provides all
users with access to all of the data. This system minimizes
the following problems:
• Data redundancy: The same data are stored in multiple
locations.
• Data isolation: Applications cannot access data associated
with other applications.
• Data inconsistency: Various copies of the data do not
agree.
12/17/2024 25
In addition, database systems maximize the following:
• Data security: Because data are “put in one place” in databases,
there is a risk of losing a lot of data at once. Therefore, databases
have extremely high security measures in place to minimize
mistakes and deter attacks.
• Data integrity: Data meet certain constraints; for example, there
are no alphabetic characters in a Social Security number field.
• Data independence: Applications and data are independent of
one another; that is, applications and data are not linked to each
other, so all applications are able to access the same data.

12/17/2024 26
12/17/2024 27
The Data Hierarchy
Data are organized in a hierarchy that begins with bits and proceeds all the
way to databases (see Figure).
• A bit (binary digit) represents the smallest unit of data a computer can
process.
• The term binary means that a bit can consist only of a 0 or a 1.
• A group of eight bits, called a byte, represents a single character. A byte
can be a letter, a number, or a symbol.

12/17/2024 28
• A logical grouping of characters into a word, a small group
of words, or an identification number is called a field.
• For example, a student’s name in a university’s computer
files would appear in the “name” field, and her or his
Social Security number would appear in the “Social
Security number” field.
• Fields can also contain data other than text and numbers.
• They can contain an image, or any other type of
multimedia.
12/17/2024 29
• A logical grouping of related fields, such as the student’s name, the
courses taken, the date, and the grade, comprises a record.
• A logical grouping of related records is called a data file or a
table.
• For example, a grouping of the records from a particular course,
consisting of course number, professor, and students’ grades,
would constitute a data fi le for that course.
• Continuing up the hierarchy, a logical grouping of related files
constitutes a database.
• Using the same example, the student course fi le could be grouped
with
12/17/2024
files on students’ personal histories and financial backgrounds
30
12/17/2024 31
Database Management Systems
• A database management system (DBMS) is a set of programs that
provide users with tools to add, delete, access, modify, and analyze
data stored in a single location.
• An organization can access the data by using query and reporting tools
that are part of the DBMS or by using application programs specifically
written to perform this function.
• DBMSs also provide the mechanisms for maintaining the integrity of
stored data, managing security and user access, and recovering
information if the system fails.
• Because databases and DBMSs are essential to all areas of business,
they must be carefully managed.
12/17/2024 32
Describing Data Warehouses and Data Marts
• In general, data warehouses and data marts support business
intelligence (BI) applications.
• Business intelligence is a broad category of applications,
technologies, and processes for gathering, storing, accessing, and
analyzing data to help business users make better decisions.
• A data warehouse is a repository of historical data that are
organized by subject to support decision makers in the
organization. Because data warehouses are so expensive, they are
used primarily by large companies.

12/17/2024 33
• A data mart is a low-cost, scaled-down version of a data
warehouse that is designed for the end-user needs in a
strategic business unit (SBU) or an individual department.
• Data marts can be implemented more quickly than data
warehouses, often in less than 90 days.
• Further, they support local rather than central control by
conferring power on the user group. Typically, groups that
need a single or a few BI applications require only a data
mart, rather than a data warehouse.

12/17/2024 34
The basic characteristics of data warehouses and data marts include
the following:
Organized by business dimension or subject.
• Data are organized by subject—for example, by customer, vendor,
product, price level, and region.
• This arrangement differs from transactional systems, where data are
organized by business process, such as order entry, inventory control,
and accounts receivable.

12/17/2024 35
Use online analytical processing.
• Typically, organizational databases are oriented toward handling
transactions.
• That is, databases use online transaction processing (OLTP), where
business transactions are processed online as soon as they occur.
• The objectives are speed and efficiency, which are critical to a
successful Internet-based business operation.
• Data warehouses and data marts, which are designed to support
decision makers but not OLTP, use online analytical processing.
Online analytical processing (OLAP) involves the analysis of
accumulated data by end users.
12/17/2024 36
• Integrated.
• Data are collected from multiple systems and then
integrated around subjects.
• For example, customer data may be extracted from internal
(and external) systems and then integrated around a
customer identifier, thereby creating a comprehensive view
of the customer.

12/17/2024 37
Time variant.
• Data warehouses and data marts maintain historical
data (i.e., data that include time as a variable).
• Unlike transactional systems, which maintain only recent
data (such as for the last day, week, or month), a
warehouse or mart may store years of data.
• Organizations utilize historical data to detect deviations,
trends, and long-term relationships.

12/17/2024 38
• Nonvolatile.
• Data warehouses and data marts are nonvolatile—that
is, users cannot change or update the data.
• Therefore the warehouse or mart reflects history,
which, as we just saw, is critical for identifying and
analyzing trends.
• Warehouses and marts are updated, but through IT-
controlled load processes rather than by users.

12/17/2024 39
• Multidimensional.
• Typically the data warehouse or mart uses a multidimensional
data structure.
• Recall that relational databases store data in two-dimensional
tables.
• In contrast, data warehouses and marts store data in more than
two dimensions. For this reason, the data are said to be stored
in a multidimensional structure.

12/17/2024 40
A Generic Data Warehouse Environment
The environment for data warehouses and marts includes the
following:
• Source systems that provide data to the warehouse or mart
• Data-integration technology and processes that prepare the data
for use.
• Different architectures for storing data in an organization’s data
warehouse or data marts
• Different tools and applications for the variety of users.
• Metadata, data-quality, and governance processes that ensure
that the warehouse or mart meets its purposes
12/17/2024 41
12/17/2024 42
Knowledge Management
• Knowledge management (KM) is a process that helps organizations
manipulate important knowledge that comprises part of the organization’s
memory, usually in an unstructured format. For an organization to be
successful, knowledge, as a form of capital, must exist in a format that can
be exchanged among persons. In addition, it must be able to grow.
Knowledge. In the information technology context, knowledge is distinct
from data and information. Data are a collection of facts, measurements, and
statistics; information is organized or processed data that are timely and
accurate. Knowledge is information that is contextual, relevant, and useful.
Simply put, knowledge is information in action.
Intellectual capital (or intellectual assets) is another term for
knowledge.
12/17/2024 43
Explicit and Tacit Knowledge.
• Explicit knowledge deals with more objective, rational, and technical
knowledge.
• In an organization, explicit knowledge consists of the policies, procedural
guides, reports, products, strategies, goals, core competencies, and IT
infrastructure of the enterprise.
• In other words, explicit knowledge is the knowledge that has been
codified (documented) in a form that can be distributed to others or
transformed into a process or a strategy.
• A description of how to process a job application that is documented in a
firm’s human resources policy manual is an example of explicit
knowledge.
12/17/2024 44
• In contrast, tacit knowledge is the cumulative store of subjective
or experiential learning.
• In an organization, tacit knowledge consists of an organization’s
experiences, insights, expertise, know-how, trade secrets, skill sets,
understanding, and learning.
• It also includes the organizational culture, which reflects the past
and present experiences of the organization’s people and
processes, as well as the organization’s prevailing values.

12/17/2024 45
• Tacit knowledge is generally imprecise and costly to transfer. It is also
highly personal.
• Finally, because it is unstructured, it is difficult to formalize or codify, in
contrast to explicit knowledge.
• A salesperson who has worked with particular customers over time and
has come to know their needs quite well would possess extensive tacit
knowledge.
• This knowledge is typically not recorded. In fact, it might be difficult for
the salesperson to put into writing, even if he or she were willing to
share it.

12/17/2024 46
Knowledge Management Systems
• The goal of knowledge management is to help an organization make
the most productive use of the knowledge it has accumulated.
Historically, management information systems have focused on
capturing, storing, managing, and reporting explicit knowledge.
• Organizations now realize they need to integrate explicit and tacit
knowledge into formal information systems.

12/17/2024 47
• Knowledge management systems (KMSs) refer to the use of
modern information technologies—the Internet, intranets,
extranets, databases—to systematize, enhance, and expedite
intrafirm and interfirm knowledge management.
• KMSs are intended to help an organization cope with turnover,
rapid change, and downsizing by making the expertise of the
organization’s human capital widely accessible.

12/17/2024 48
• Organizations can realize many benefits with KMSs.
• Most importantly, they make best practices, the most effective
and efficient ways of doing things, readily available to a wide range
of employees. Enhanced access to best-practice knowledge
improves overall organizational performance.
• For example, account managers can now make available their tacit
knowledge about how best to manage large accounts. The
organization can then utilize this knowledge when it trains new
account managers. Other benefits include improved customer
service, more efficient product development, and improved
employee morale and retention.
12/17/2024 49
The KMS Cycle
A functioning KMS follows a cycle
that consists of six steps (see
Figure ). The reason the system is
cyclical is that knowledge is
dynamically refined over time. The
knowledge in an effective KMS is
never finalized because the
environment changes over time and
knowledge must be updated to
reflect these changes.
12/17/2024 50
The cycle works as follows:
1. Create knowledge. Knowledge is created as people determine
new ways of doing things or develop know-how.
Sometimes external knowledge is brought in.
2. Capture knowledge. New knowledge must be identified as
valuable and be represented in a reasonable way.
3. Refine knowledge. New knowledge must be placed in context so
that it is actionable. This is where tacit qualities (human insights)
must be captured along with explicit facts.

12/17/2024 51
4. Store knowledge. Useful knowledge must then be stored
in a reasonable format in a knowledge repository so that others
in the organization can access it.
5. Manage knowledge. Like a library, the knowledge must be
kept current. It must be reviewed regularly to verify that it is
relevant and accurate.
6. Disseminate knowledge. Knowledge must be made
available in a useful format to anyone in the organization who
needs it, anywhere and anytime.

12/17/2024 52

You might also like