0% found this document useful (0 votes)
29 views54 pages

Text Analytics and Mining Explained

The document provides an overview of Text Analytics and Text Mining, highlighting their importance in processing unstructured text data, which constitutes 85% of corporate data. It distinguishes between the two concepts, with Text Analytics being a broader process that includes Information Retrieval and Text Mining, which focuses on discovering patterns in text. The document also discusses the applications, benefits, and technologies involved in Text Mining, emphasizing its role in various domains such as law, finance, and marketing.

Uploaded by

drlngismail67
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
29 views54 pages

Text Analytics and Mining Explained

The document provides an overview of Text Analytics and Text Mining, highlighting their importance in processing unstructured text data, which constitutes 85% of corporate data. It distinguishes between the two concepts, with Text Analytics being a broader process that includes Information Retrieval and Text Mining, which focuses on discovering patterns in text. The document also discusses the applications, benefits, and technologies involved in Text Mining, emphasizing its role in various domains such as law, finance, and marketing.

Uploaded by

drlngismail67
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

BUSINESS ANALYTICS BAD714B

MODULE-5

Text Analytics and Text Mining Overview

Introduction: The Importance of Text Data


We live in an information age, where the amount of data generated, collected, and stored is
growing at an exponential rate.
Most of this data is unstructured text, such as:
 Emails
 Reports
 Web pages
 Social media posts
 Product reviews
 Research articles
A study by Merrill Lynch and Gartner found that:
 85% of all corporate data exists in unstructured text form
 The volume of such data doubles every 18 months
Because knowledge is power in business, and knowledge comes from analyzing data,
companies that can effectively process and understand their text data gain a strong
competitive advantage. This is where Text Analytics and Text Mining become essential.

What Are Text Analytics and Text Mining?

Both terms deal with transforming unstructured text into structured, actionable
information using Natural Language Processing (NLP) and analytics techniques.
There are subtle differences:
Term Focus Scope
A broader concept involving the full Includes information retrieval,
Text
process of retrieving, extracting, and information extraction, data mining,
Analytics
analyzing text data and web mining

Text A specific process of discovering new Primarily focused on pattern discovery


Mining and useful knowledge from textual data within text

Relationship between the two:


Text Analytics=Information Retrieval+Information Extraction+Data Mining+Web Mining
or simply,
Text Analytics=Information Retrieval+Text Mining
So, text analytics is an umbrella term, and text mining is one of its core components.

Difference Between Text Analytics and Text Mining


Aspect Text Analytics Text Mining
Broad process of extracting, analyzing, and Focused on discovering patterns and
Definition
visualizing information from text data insights from textual sources
Usage More common in business contexts More common in academic research
Goal Turn text into structured data and insights Discover new and useful knowledge

Dept of AI&DS, SIET Page 1


BUSINESS ANALYTICS BAD714B

Aspect Text Analytics Text Mining


Covers multiple subfields (IR, IE, Web Mainly pattern recognition and
Scope
Mining) knowledge discovery

In practice, the two terms are often used interchangeably.

Text Mining

Text Mining (also known as Text Data Mining or Knowledge Discovery in Textual
Databases) is a semi-automated process that:

 Extracts patterns and useful information from large volumes of unstructured text
 Converts unstructured text into structured data for analysis

It is conceptually similar to Data Mining, but differs in the type of data processed:
Aspect Data Mining Text Mining
Structured (e.g., databases, Unstructured (e.g., documents, articles,
Data Type
spreadsheets) emails)
Input
Numeric/categorical values Text, XML, PDFs, Word files
Format
Impose structure on text and extract
Process Discover patterns in structured data
patterns
Classification, clustering, NLP, parsing, information extraction,
Techniques
association rules classification

The Two Main Steps of Text Mining:

1. Text Preprocessing and Structuring


o Convert unstructured text into structured form
o Tasks include:
 Tokenization (splitting text into words)
 Stop-word removal (removing common words like “the”, “and”)
 Stemming/Lemmatization (reducing words to their base form)
 Part-of-speech tagging
 Creating document-term matrices
2. Pattern Discovery and Knowledge Extraction
o Apply data mining algorithms (e.g., clustering, classification, association) on
structured text data
o Discover trends, patterns, and relationships

Benefits and Application Areas of Text Mining

Text mining has become crucial in areas where large text volumes are produced daily:
Domain Examples of Text Mining Use
Law Analyze court orders, judgments, and legal precedents
Academia Identify emerging research topics from scientific articles
Finance Analyze annual/quarterly financial reports
Medicine Extract data from patient discharge summaries or medical notes

Dept of AI&DS, SIET Page 2


BUSINESS ANALYTICS BAD714B

Domain Examples of Text Mining Use


Biology Discover molecular interactions from research papers
Technology Examine patents for innovation trends
Marketing&Customer
Analyze feedback, reviews, complaints for product improvement
Service

Example:

Customer feedback and warranty claims can reveal product flaws and customer sentiment.
By analyzing such text, businesses can:

 Improve product design


 Enhance customer satisfaction
 Predict trends in consumer behavior

Another example is email processing:

 Spam detection (classify unwanted emails)


 Automatic prioritization (based on importance)
 Auto-response generation (using pattern recognition)

Applications of Text Mining

a. Information Extraction (IE)

 Identifies key phrases, entities, and relationships within text.


 Uses pattern matching to find predefined objects (like names, dates, places).
 Example: Extracting names of companies from financial reports.

b. Topic Tracking

 Predicts new documents that might interest a user based on their reading history or
profile.
 Example: Recommending news articles similar to those already read.

c. Summarization

 Automatically generates concise summaries of long documents.


 Saves time for users by presenting the essential content.
 Example: Executive summaries of research papers or news.

d. Categorization (Text Classification)

 Assigns documents to predefined categories based on content.


 Example: Sorting news into “Politics,” “Sports,” “Technology,” etc.

e. Clustering

 Groups similar documents without predefined labels.


 Example: Grouping customer complaints with similar themes.

Dept of AI&DS, SIET Page 3


BUSINESS ANALYTICS BAD714B

f. Concept Linking

 Connects related documents through shared concepts.


 Helps discover hidden relationships not found by traditional keyword searches.
 Example: Linking “global warming” articles with “carbon emissions.”

g. Question Answering

 Finds the best possible answer to a question from text databases.


 Uses knowledge-driven pattern matching and NLP.
 Example: Chatbots, search engines (like Google’s “featured snippets”).

Technology Insights

To perform these tasks, Text Mining systems rely on:

 Natural Language Processing (NLP) — to understand and process human language


 Machine Learning — to learn from text patterns
 Information Retrieval (IR) — to locate relevant documents
 Linguistics — to analyze grammatical and semantic structure

Together, these fields form the foundation of text analytics and mining.

Real-World Impact

Organizations using text analytics and mining can:

 Enhance decision-making by leveraging hidden insights


 Improve customer experience by understanding opinions
 Increase efficiency by automating data analysis
 Achieve competitive advantage by faster trend detection

Dept of AI&DS, SIET Page 4


BUSINESS ANALYTICS BAD714B

FIGURE: Text Analytics, Related Application Areas, and Enabling Disciplines.

Application Case 5.1: Insurance Group Strengthens Risk Management with Text
Mining Solution
When asked for the biggest challenge facing the Czech automobile insurance industry, Peter
Jedlicˇka, PhD, doesn’t hesitate. “Bodily injury claims are growing disproportionately
compared with vehicle damage claims,” says Jedlicˇka, team leader of actuarial services for
the Czech Insurers’ Bureau (CIB). CIB is a professional organization of insurance companies
in the Czech Republic that handles uninsured, international, and untraced claims for what’s
known as motor third-party liability. “Bodily injury damages now represent about 45% of the
claims made against our members, and that proportion will continue to increase because of
recent legislative changes.” One of the difficulties that bodily injury claims pose for insurers
is that the extent of an injury is not always predictable in the immediate aftermath of a
vehicle accident. Injuries that were not at first obvious may become acute later, and
apparently minor injuries can turn into chronic conditions. The earlier that insurance
companies can accurately estimate their liability for medical damages, the more precisely
they can manage their risk and consolidate their resources. However, because the needed
information is contained in unstructured documents such as accident reports and witness
statements, it is extremely time consuming for individual employees to perform the needed
analysis. To expand and automate the analysis of unstructured accident reports, witness
statements, and claim narratives, CIB deployed a data analysis solution based on Dell
Statistica Data Miner and the Statistica Text Miner extension. Statistica Data Miner offers a
set of intuitive, user-friendly tools that are accessible even to nonanalysts. Application Case
5.1 Insurance Group Strengthens Risk Management with Text Mining Solution The solution
reads and writes data from virtually all standard file formats and offers strong, sophisticated
data cleaning tools. It also supports even novice users with query wizards, called Data Mining
Recipes, that help them arrive at the answers they need more quickly. With the Statistica Text
Miner extension, users have access to extraction and selection tools that can be used to index,
classify, and cluster information from large collections of unstructured text data, such as the

Dept of AI&DS, SIET Page 5


BUSINESS ANALYTICS BAD714B

narratives of insurance claims. In addition to using the Statistica solution to make predictions
about future medical damage claims, CIB can also use it to find patterns that indicate
attempted fraud or to identify needed road safety improvements.
Improves Accuracy of Liability Estimates
Jedlicˇka expects the Statistica solution to greatly improve the ability of CIB to predict the
total medical claims that might arise from a given accident. “The Statistica solution’s data
mining and text mining capabilities are already helping us expose additional risk
characteristics, thus making it possible to predict serious medical claims in earlier stages of
the investigation,” he says. “With the Statistica solution, we can make much more accurate
estimates of total damages and plan accordingly.”
Expands Service Offerings to Members
Jedlicˇka is also pleased that the Statistica solution helps CIB offer additional services to its
member companies. “We are in a data-driven business,” he says. “With Statistica, we can
provide our members with detailed analyses of claims and market trends. Statistica also helps
us provide even stronger recommendations concerning claims reserves.”
Intuitive for Business Users
The intuitive Statistica tools are accessible by even nontechnical users. “The outputs of our
Statistica analyses are easy to understand for business users,” says Jedlicˇka. “Our business
users also find that the analysis results are in line with their own experience and
recommendations, so they readily see the value in the Statistica solution.”
Questions for Discussion
1. How can text analytics and mining be used to keep up with changing business needs of
insurance companies?
2. What were the challenges, the proposed solution, and the obtained results?
3. Can you think of other uses of text analytics and text mining for insurance companies?

Natural Language Processing (NLP)

Introduction
Natural Language Processing (NLP) is a crucial subfield of Artificial Intelligence (AI)
and Computational Linguistics.
It focuses on enabling computers to understand, interpret, and generate human language
in a way that is both meaningful and useful.

In the context of Text Mining and Text Analytics, NLP helps transform unstructured
textual data (words, sentences, paragraphs) into structured formats (numeric or symbolic
data) that can be easily analyzed by algorithms.

From Bag-of-Words to NLP

a. Early Methods: The Bag-of-Words Model

One of the earliest methods used in text mining was the Bag-of-Words (BoW) approach.
In this method:

 A document (sentence, paragraph, or full text) is represented as a collection (or


“bag”) of words.
 It ignores grammar, word order, and sentence structure.
 Only the frequency or occurrence of words is considered.

Example:

Dept of AI&DS, SIET Page 6


BUSINESS ANALYTICS BAD714B

Let’s take two sentences:

1. “Dogs chase cats.”


2. “Cats chase dogs.”

Under the bag-of-words model, both sentences are represented by the same words {dogs,
chase, cats}, even though their meaning is different.

Hence, BoW loses semantic meaning.

b. Use of Bag-of-Words in Applications

Even today, simple tasks such as spam filtering use the BoW model effectively.

Example: Spam Email Filtering

 Two bags are created:


o One for spam emails
o One for legitimate (ham) emails
 If a new email contains more words found in the spam bag (e.g., Viagra, buy, stock,
offer), it is classified as spam.

Although this method works for basic categorization, it does not truly understand the
meaning or context of language.

c. Limitation of Bag-of-Words

While BoW can capture word frequency, it fails to:

 Understand syntax (grammar rules)


 Capture semantics (meaning of words in context)
 Handle polysemy (multiple meanings of the same word)
 Recognize relationships between words

For example, in medical text classification, researchers found that using BoW representation
on millions of abstracts from MEDLINE yielded poor results — nearly equivalent to random
guessing — because the model lacked contextual and semantic understanding.

Thus, researchers moved toward more advanced linguistic techniques, giving rise to NLP.

What is NLP?

Definition:
Natural Language Processing is a computational approach to analyzing and
understanding human language, converting textual information into structured data that
computers can process.

It aims to go beyond simple word counting and incorporate:

 Syntax (grammatical structure)


 Semantics (meaning)

Dept of AI&DS, SIET Page 7


BUSINESS ANALYTICS BAD714B

 Context (situational relevance)

The goal of NLP is to achieve a “true understanding” of language — enabling machines not
just to process text, but to interpret its meaning.

Challenges in NLP

Understanding human language is extremely difficult for computers because natural


language is:
 Ambiguous
 Context-dependent
 Irregular
 Full of idioms, sarcasm, and metaphors

There are some major challenges faced in NLP:


Challenge Description Example
Determining whether a word is a noun, The word “book” can mean a
1. Part-of-Speech
verb, adjective, etc., depending on its noun (“read a book”) or a verb
Tagging
usage. (“book a ticket”).
Identifying boundaries between words Segmenting continuous text
2. Text and sentences, especially in languages
correctly in “我喜欢吃苹果”
Segmentation without spaces (e.g., Chinese, Japanese,
(I like to eat apples).
Thai).
Choosing the correct meaning of a word “Bank” can mean a financial
3. Word Sense
that has multiple meanings, based on institution or the side of a
Disambiguation
context. river.
“I saw the man with the
4. Syntactic The same sentence structure can have
telescope.” → Who had the
Ambiguity multiple interpretations.
telescope?
5. Imperfect or Handling errors, slang, typos, accents, “Wanna go?” instead of “Want
Irregular Input or regional variations in text or speech. to go?”
Understanding the intention behind a “Can you pass the salt?” is a
6. Speech Acts
statement, not just its structure. request, not a question.

NLP in Knowledge Extraction

It has been a long-standing dream of AI researchers to enable computers to read and


acquire knowledge automatically from text.

Stanford NLP Example:

 Stanford researchers applied learning algorithms to parsed text to automatically:


o Identify concepts (key ideas)
o Detect relationships between concepts
 They enriched WordNet, a large lexical database of English words, by automatically
adding new relationships learned from text.

This greatly expanded the usability of WordNet while reducing manual effort.

Dept of AI&DS, SIET Page 8


BUSINESS ANALYTICS BAD714B

WordNet and Its Role in NLP

WordNet is a hand-coded database of English words and their:


 Synonyms (synsets)
 Definitions
 Semantic relationships (e.g., “is-a,” “part-of,” etc.)

It serves as a knowledge base for many NLP applications such as:

 Semantic search
 Text classification
 Sentiment analysis

By automating the addition of knowledge to WordNet through NLP algorithms, we can


make it more comprehensive and cost-effective.

Applications of NLP

NLP powers a wide range of real-world applications across industries.


Here are some major NLP tasks and their functions:
Task Description
Automatically answers user questions in natural language
Question Answering (QA)
(e.g., Chatbots, Google Assistant).
Generates concise summaries of long documents or news
Automatic Summarization
articles.
Natural Language Generation Converts structured data into human-readable text (e.g.,
(NLG) report generation).
Natural Language Converts text into formal, structured representations
Understanding (NLU) understandable by computers.
Automatically translates text from one language to another
Machine Translation
(e.g., Google Translate).
Assists users in reading foreign languages with
Foreign Language Reading
pronunciation aids.
Suggests grammatical corrections for non-native writers
Foreign Language Writing
(e.g., Grammarly).
Speech Recognition Converts spoken language into text (e.g., voice typing, Siri).
Converts written text into natural-sounding speech (e.g.,
Text-to-Speech (TTS)
accessibility tools).
Detects and corrects grammar, spelling, and punctuation
Text Proofing
errors.
Optical Character
Converts scanned images of text into machine-readable text.
Recognition (OCR)

NLP in Sentiment Analysis (Example Application)

One powerful application of NLP is Sentiment Analysis — a technique that determines


positive, negative, or neutral opinions expressed in text.

Dept of AI&DS, SIET Page 9


BUSINESS ANALYTICS BAD714B

Example:

In Customer Relationship Management (CRM):

 NLP is used to analyze customer reviews, feedback, and social media posts.
 It helps identify customer emotions toward products or services.
 Businesses can detect dissatisfaction early, improve services, and target marketing
efforts effectively.

Role of NLP in Text Mining

The success of Text Mining largely depends on advancements in NLP.


Aspect Text Mining NLP’s Contribution
Input Data Unstructured text Converts to structured data
Goal Discover patterns, knowledge Provides language understanding
Data mining, ML, pattern Tokenization, parsing, tagging, semantic
Techniques
recognition analysis

In simple terms:

Text Mining=NLP+Data Mining

That means NLP provides language structure and meaning, and data mining algorithms
discover patterns and insights from that structured data.

Application Case 5.2: AMC Networks Is Using Analytics to Capture New Viewers,
Predict Ratings, and Add Value for Advertisers in a Multichannel World
Over the past 10 years, the cable television sector in the United States has enjoyed a period of
growth that has enabled unprecedented creativity in the creation of high-quality content.
AMC Networks has been at the forefront of this new golden age of television, producing a
string of successful, critically acclaimed shows such as Breaking Bad, Mad Men, and The
Walking Dead. Dedicated to producing quality programming and movie content for more
than 30 years, AMC Networks Inc. owns and operates several of the most popular and award-
winning brands in cable television, producing and delivering distinctive, compelling, and
culturally relevant content that engages audiences across multiple platforms.
Getting Ahead of the Game
Despite its success, the company has no plans to rest on its laurels. As Vitaly Tsivin, SVP
Business Intelligence, explains: “We have no interest in standing still. Although a large
percentage of our business is still linear cable TV, we need to appeal to a new generation of
millennials who consume content in very different ways. “TV has evolved into a
multichannel, multistream business, and cable networks need to get smarter about how they
market to and connect with audiences across all of those streams. Relying on traditional
ratings data and third-party analytics providers is going to be a losing strategy: you need to
take ownership of your data, and use it to get a richer picture of who your viewers are, what
they want, and how you can keep their attention in an increasingly crowded entertainment
marketplace.”
Zoning in on the Viewer
The challenge is that there is just so much information available—hundreds of billions of
rows of data from industry data providers such as Nielsen and comScore, from channels such
as AMC’s TV Everywhere live Web streaming and video on-demand service, from retail

Dept of AI&DS, SIET Page 10


BUSINESS ANALYTICS BAD714B

partners such as iTunes and Amazon, and from third-party online video services such as
Netflix and Hulu. “We can’t rely on high-level summaries; we need to be able to analyze
both structured and unstructured data, minute-by-minute and viewer-byviewer,” says Vitaly
Tsivin. “We need to know who’s watching and why—and we need to know it quickly so that
we can decide, for example, whether to run an ad or a promo in a particular slot during
tomorrow night’s episode of Mad Men.” AMC decided it needed to develop an
industryleading analytics capability in-house—and focused on delivering this capability as
quickly as possible. Instead of conducting a prolonged and expensive vendor and product
selection process, AMC decided to leverage its existing relationship with IBM as its trusted
strategic technology partner. The time and money traditionally spent on procurement were
instead invested in realizing the solution—accelerating AMC’s progress on its analytics
roadmap by at least 6 months.
Empowering the Research Department
In the past, AMC’s research team spent a large portion of its time processing data. Today,
thanks to its new analytics tools, it is able to focus most of its energy on gaining actionable
insights. “By investing in big data analytics technology from IBM, we’ve been able to
increase the pace and detail of our research an order of magnitude,” says Vitaly Tsivin.
“Analyses that used to take days and weeks are now possible in minutes, or even seconds.
“Bringing analytics in-house will provide major ongoing cost-savings. Instead of paying
hundreds of thousands of dollars to external vendors when we need some analysis done, we
can do it ourselves— more quickly, more accurately, and much more costeffectively. We’re
expecting to see a rapid return on investment. “As more sources of potential insight become
available and analytics becomes more strategic to the business, an in-house approach is really
the only viable way forward for any network that truly wants to gain competitive advantage
from its data.”

FIGURE: A Web-Based Dashboard Used by AMC Networks.


Driving Decisions with Data
Many of the results delivered by this new analytics capability demonstrate a real
transformation in the way AMC operates. For example, the company’s business intelligence
department has been able to create sophisticated statistical models that help the company
refine its marketing strategies and make smarter decisions about how intensively it should
promote each show. With deeper insight into viewership, AMC’s direct marketing campaigns

Dept of AI&DS, SIET Page 11


BUSINESS ANALYTICS BAD714B

are also much more successful. In one recent example, intelligent segmentation and lookalike
modeling helped the company target new and existing viewers so effectively that AMC video
on-demand transactions were higher than would be expected otherwise. This newfound
ability to reach out to new viewers based on their individual needs and preferences is not just
valuable for AMC—it also has huge potential value for the company’s advertising partners.
AMC is currently working on providing access to its rich data sets and analytics tools as a
service for advertisers, helping them fine-tune their campaigns to appeal to ever-larger
audiences across both linear and digital channels. Vitaly Tsivin concludes: “Now that we can
really harness the value of big data, we can build a much more attractive proposition for both
consumers and advertisers—creating even better content, marketing it more effectively, and
helping it reach a wider audience by taking full advantage of our multichannel capabilities.”

Questions for Discussion


1. What are the common challenges broadcasting companies are facing nowadays? How can
analytics help to alleviate these challenges?
2. How did AMC leverage analytics to enhance their business performance?
3. What were the types of text analytics and text mini solutions developed by AMC
networks? Can you think of other potential uses of text mining applications in the
broadcasting industry?

Text Mining Applications


Text mining refers to extracting useful and meaningful information from unstructured text
data, such as emails, social media posts, research papers, or customer feedback.
As organizations collect more unstructured data, text mining tools help convert this data into
actionable insights.
Overview
 Purpose: To extract patterns, trends, and relationships from large collections of
textual information.
 Data Sources: Documents, emails, reports, call center notes, social media, scientific
literature, etc.
 Goal: Turn unstructured text into structured information for decision-making.

Dept of AI&DS, SIET Page 12


BUSINESS ANALYTICS BAD714B

Major Application Areas


Text mining has broad applications across business, security, biomedical, and academic
fields.
Let’s explore each in detail:
A. Marketing Applications
i. Customer Relationship Management (CRM)
 Text mining helps companies analyze unstructured data from call centers, emails,
or chat logs.
 It identifies customer opinions, complaints, and satisfaction levels.
 For example, analyzing call center transcripts can reveal patterns in customer
sentiment toward a product or service.
ii. Sentiment Analysis
 Text from blogs, social media, and product reviews can be mined to understand
public perception.
 Businesses use this information to:
o Improve products and services.
o Identify customer pain points.
o Design marketing strategies to increase cross-selling and up-selling
opportunities.
iii. Customer Retention and Churn Prediction
 Coussement & Van den Poel (2009) showed that text mining can predict customer
churn (attrition) — customers likely to leave a company.
 By combining textual and structured data, companies can identify at-risk customers
and design retention strategies.
iv. Product Attribute Discovery
 Ghani et al. (2006) developed a text-mining system to automatically infer product
attributes (e.g., color, size, brand) from web descriptions.
 This helps in:
o Product recommendation systems.
o Demand forecasting.
o Assortment optimization (choosing what to stock).
o Supplier selection.
 Uses supervised and semi-supervised learning to extract product features with
minimal manual work.
B. Security Applications
Text mining plays a vital role in national security and law enforcement by analyzing huge
volumes of communication data.
i. Surveillance Systems
 The ECHELON system (rumored) intercepts and analyzes global communications
such as emails, calls, and faxes for intelligence purposes.
 Uses text mining to detect keywords, topics, and threats.
ii. Law Enforcement – OASIS System
 EUROPOL developed the OASIS (Overall Analysis System for Intelligence
Support) in 2007.
 Integrates structured and unstructured data using data and text mining for tracking
organized crime across countries.
 Improves international crime investigation and intelligence sharing.
iii. U.S. Intelligence System
 The FBI and CIA, under the Department of Homeland Security, developed a joint
supercomputer system to integrate data and text mining for intelligence.

Dept of AI&DS, SIET Page 13


BUSINESS ANALYTICS BAD714B

 Before this, both had separate databases; now, integration enables cross-agency data
analysis to detect threats efficiently.
iv. Deception Detection
 Fuller, Biros, & Delen (2008) developed text mining models to detect deceptive vs.
truthful statements from criminal text data.
 Achieved around 70% accuracy using only textual cues (no visual or vocal data).
 Compared to polygraphs, this approach is non-intrusive and applicable to large text
corpora (like police interviews or online statements).
C. Biomedical Applications
The biomedical field produces vast amounts of literature and experimental data (e.g., gene,
protein, disease information). Text mining helps interpret and connect these findings.
i. Literature Analysis and Knowledge Discovery
 Biomedical research generates enormous data through experiments like:
o DNA microarrays,
o SAGE (Serial Analysis of Gene Expression),
o Proteomics (mass spectrometry).
 Scientists use text mining to integrate new experimental results with existing
biomedical literature for better understanding of biological entities.
ii. Protein Location Prediction
 Shatkay et al. (2007) developed a text-mining system combining sequence-based
and text-based features to predict protein locations within cells.
 Knowing where a protein resides helps understand its biological role and drug target
potential.
 Their hybrid approach outperformed earlier models.
iii. Disease–Gene Relationship Extraction
 Chun et al. (2006) built a system that extracts relationships between genes and
diseases from MEDLINE.
 Used:
o A dictionary-based approach for gene/disease names.
o A machine learning–based Named Entity Recognition (NER) system to
filter false positives.
 Result: 26.7% improvement in precision, with only a slight drop in recall.
iv. Gene–Protein Relationship Discovery
 Nakov et al. (2005) illustrated how multi-level text analysis can uncover gene–
protein interactions.
 The process involves:
1. Tokenization (breaking text into words),
2. Part-of-Speech tagging,
3. Parsing and matching terms to biological ontologies (structured domain
knowledge).
 This helps decode complex relationships in projects like the Human Genome
Project.

Dept of AI&DS, SIET Page 14


BUSINESS ANALYTICS BAD714B

FIGURE: Multilevel Analysis of Text for Gene/Protein Interaction Identification.

D. Academic Applications
Text mining supports knowledge organization, indexing, and retrieval in academic and
publishing domains.
i. Scientific Publishing
 Publishers maintain large digital databases needing automatic indexing for easier
retrieval.
 Initiatives include:
o Nature’s Open Text Mining Interface – allows semantic querying.
o NIH Journal Publishing Document Type Definition – standardizes
document structures for automated mining.
ii. Academic Research Centers
 National Centre for Text Mining (NaCTeM) – collaboration between the
University of Manchester and University of Liverpool.
o Initially focused on biomedical sciences, now extended to social sciences.
o Provides tools, training, and customized solutions for academic text mining.
 BioText Project (UC Berkeley) – supports bioscience researchers in text analysis for
biomedical discovery.
iii. Importance in Academia
 Helps researchers discover hidden connections, trends, and hypotheses.
 Enables semantic search, where computers understand meaning, not just keywords.

Application Case 5.3: Mining for Lies


Driven by advancements in Web-based information technologies and increasing
globalization, computermediated communication continues to filter into everyday life,
bringing with it new venues for deception. The volume of text-based chat, instant messaging,
text messaging, and text generated by online communities of practice is increasing rapidly.
Even e-mail continues to grow in use. With the massive growth of text-based communication,
the potential for people to deceive others through computermediated communication has also
grown, and such deception can have disastrous results. Unfortunately, in general, humans
tend to perform poorly at deception-detection tasks. This phenomenon is exacerbated in text-
based communications. A large part of the research on deception detection (also known as

Dept of AI&DS, SIET Page 15


BUSINESS ANALYTICS BAD714B

credibility assessment) has involved face-to-face meetings and interviews. Yet, with the
growth of text-based communication, text-based deception-detection techniques are essential.
Techniques for successfully detecting deception—that is, lies—have wide applicability. Law
enforcement can use decision support tools and techniques to investigate crimes, conduct
security screening in airports, and monitor communications of suspected terrorists. Human
resources professionals might use deception-detection tools to screen applicants. These tools
and techniques also have the potential to screen e-mails to uncover fraud or other
wrongdoings committed by corporate officers. Although some people believe that they can
readily identify those who are not being truthful, a summary of deception research showed
that, on average, people are only 54% accurate in making veracity determinations (Bond &
DePaulo, 2006). This figure may actually be worse when humans try to detect deception in
text. Using a combination of text mining and data mining techniques, Fuller et al. (2008)
analyzed person-of-interest statements completed by people involved in crimes on military
bases. In these statements, suspects and witnesses are required to write their recollection of
the event in their own words. Military law enforcement personnel searched archival data for
statements that they could conclusively identify as being truthful or deceptive. These
decisions were made on the basis of corroborating evidence and case resolution. Once labeled
as truthful or deceptive, the law enforcement personnel removed identifying information and
gave the statements to the research team. In total, 371 usable statements were received for
analysis. The text-based deception-detection method used by Fuller et al. (2008) was based
on a process known as message feature mining, which relies on elements of data and text
mining techniques. A simplified depiction of the process is provided in Figure 5.3. First, the
researchers prepared the data for processing. The original handwritten statements had to be
transcribed into a word processing file. Second, features (i.e., cues) were identified. The
researchers identified 31 features representing categories or types of language that are
relatively independent of the text content and that can be readily analyzed by automated
means. For example, first-person pronouns such as I or me can be identified without analysis
of the surrounding text. Table 5.1 lists the categories and an example list of features used in
this study. The features were extracted from the textual statements and input into a flat file
for further processing. Using several feature-selection methods along with 10-fold cross-
validation, the researchers compared the prediction accuracy of three popular data mining
methods. Their results indicated that neural network models performed the best, with 73.46%
prediction accuracy on test data samples; decision trees performed second best, with 71.60%
accuracy; and logistic regression was last, with 65.28% accuracy. The results indicate that
automated text-based deception detection has the potential to aid those who must try to detect
lies in text and can be successfully applied to real-world data. The accuracy of these
techniques exceeded the accuracy of most other deception-detection techniques, even though
it was limited to textual cues.

Dept of AI&DS, SIET Page 16


BUSINESS ANALYTICS BAD714B

FIGURE: Text-Based Deception-Detection Process.

Table: Categories and Examples of Linguistic Features Used in Deception Detection

Questions for Discussion


1. Why is it difficult to detect deception?
2. How can text/data mining be used to detect deception in text?
3. What do you think are the main challenges for such an automated system?

Application Case 5.4: Bringing the Customer into the Quality Equation: Lenovo Uses
Analytics to Rethink Its Redesign
Lenovo was near final design on an update to the keyboard layout of one of its most popular
PCs when it spotted a small, but significant, online community of gamers who are
passionately supportive of the current keyboard design. Changing the design may have led to
a mass revolt of a large segment of Lenovo’s customer base—freelance developers and
gamers. The Corporate Analytics unit was using SAS as part of a perceptual quality project.
Crawling the web, sifting through text data for Lenovo mentions, the analysis unearthed a
previously unknown forum, where an existing customer had written a glowing six-page
review of the current design, especially the keyboard. The review attracted 2,000 comments!
“It wasn’t something we would have found in traditional preproduction design reviews,” says
Mohammed Chaara, Director of Customer Insight & VOC Analytics. It was the kind of
discovery that solidified Lenovo’s commitment to the Lenovo Early Detection (LED) system,
and the work of Chaara and his corporate analytics team. Lenovo, the largest global

Dept of AI&DS, SIET Page 17


BUSINESS ANALYTICS BAD714B

manufacturer of PCs and tablets, didn’t set out to gauge sentiment around obscure bloggers or
discover new forums. The company wanted to inform quality, product development, and
product innovation by studying data—its own and that from outside the four walls. “We’re
mainly focused on supply chain optimization, crosssell/up-sell opportunities and pricing and
packaging of services. Any improvements we make in these areas are based on listening to
the customer,” Chaara says. SAS provides the framework to “manage the crazy amount of
data” that is generated. The project’s success has traveled like wildfire within the
organization. Lenovo initially planned on about 15 users, but word of mouth has led to 300
users signing up to log in to the LED dashboard for a visual presentation on customer
sentiment, warranty, and call center analysis.
The Results Have Been Impressive
• Over 50% reduction in issue detection time.
• 10 to 15% reduction in warranty costs from out-of-norm defects.
• 30 to 50% reduction in general information calls to the contact center.
Looking at the Big Picture
Traditional methods of gauging sentiment and understanding quality have built-in
weaknesses and time lags:
• Customer surveys only surface information from customers who are willing to fill them out.
• Warranty information often comes in months after delivery of the new product.
• It can be difficult to decipher myriad causes of customer discontent and product issues.
In addition, Lenovo sells its product packaged with software it doesn’t produce, and
customers use a variety of accessories (docking stations and mouse devices) that might or
might not be Lenovo products. To compound the issue, the company operates in 165
countries and supports more than 30 languages, so the manual methods to evaluate the
commentary were inconsistent, took too much time, and couldn’t scale to the volumes of
feedback it was seeing in social media. The sentiment analysis needed to be able to sense
nuances within the native languages. (For example, Australians describe things differently
than Americans.) The analysis-driven discovery of an issue with docking stations provided
the second big win for Lenovo’s LED initiative. Customers were calling tech support to say
they were having issues with the screen, or the machine shutting down abruptly, or the
battery wasn’t charging. Similar accounts were turning up on social media posts. Sometimes,
though not always, the customer mentioned docking. It wasn’t until Lenovo used SAS to
analyze the combination of call center notes and social media posts that the word docking
was connected to the problem, helping quality engineers figure out the root cause and issue a
software update. “We were able to pick up that feedback within weeks. It used to take 60 to
90 days because we had to wait for the reports to come back from the field,” Chaara says.
Now it takes just 15 to 30 days. That reduction in detection time has driven a 10 to 15%
reduction in warranty costs for those issues. As warranty claims cost the company about $1.2
billion yearly, this is a significant savings. Although the call center information was crucial,
the social media component was what sealed the deal. “With Twitter and Facebook, people
described what they were doing at that minute, “I docked the machine and X happened.’ It’s
raw, unbiased and so powerful,” Chaara says. An unforeseen insight was found when
analyzing what customers were saying as they got their PCs up and running. Lenovo realized
its documentation to explain its products, warranties, and the like was unclear. “There is a
cost to every call center call. With the improved documentation, we’ve seen a 30 to 50%
reduction in calls coming in for general information,” Chaara said.
Winning Praise beyond the frontlines
The project has been so successful that Chaara demoed it for the CEO. The goal is to
configure a dashboard view for the C-suite. “That’s the level of thinking from our senior
executives. They believe in this,” Chaara says. In addition, Chaara’s group will be formally
measuring the success of the effort and expanding it to measure issues like customer

Dept of AI&DS, SIET Page 18


BUSINESS ANALYTICS BAD714B

experience when buying a Lenovo product. “The application of analytics has ultimately led
us to a more holistic understanding of the concept of quality. Quality isn’t just a PC working
correctly. It’s people knowing how to use it, getting quick and accurate help from the
company, getting the non-Lenovo components to work well with the hardware, and
understanding what the customers like about the existing product—rather than just
redesigning it because product designers think it’s the right thing to do. “SAS has allowed us
to get a definition of quality from the view of the customer,” Chaara says.
Questions for Discussion
1. How did Lenovo use text analytics and text mining to improve quality and design of their
products and ultimately improve customer satisfaction?
2. What were the challenges, the proposed solution, and the obtained results?

Text Mining Process


Text mining, like data mining, follows a systematic process to extract meaningful and
actionable knowledge from unstructured textual data. Because text data is unstructured, text
mining requires additional preprocessing steps compared to traditional data mining.

Overview of the Text Mining Process


A text mining process is similar to the CRISP-DM (Cross-Industry Standard Process for
Data Mining) model used in data mining.
However, text mining emphasizes data preprocessing, since text must be converted into a
structured format before analysis.
Goal:
To transform unstructured text data into structured knowledge that supports better
decision-making.

High-Level Context Diagram (Process Components)


According to Delen & Crossland (2008), a text mining process can be viewed as a system
with inputs, outputs, controls, and mechanisms.
Component Description
Unstructured and structured data (e.g., documents, web pages,
Input
emails).
Output Knowledge or insights used for decision-making.
Controls Software/hardware limitations, privacy issues, and natural language
(Constraints) complexities.
Mechanisms
Algorithms, software tools, and domain expertise.
(Resources)
The goal is to process textual data and extract patterns, relationships, and knowledge
relevant to a specific context.

Dept of AI&DS, SIET Page 19


BUSINESS ANALYTICS BAD714B

FIGURE: Context Diagram for the Text Mining Process.

Three Main Tasks of Text Mining


The text mining process consists of three consecutive tasks:
1. Establish the Corpus
2. Create the Term–Document Matrix (TDM)
3. Extract the Knowledge
Each task produces specific outputs that feed into the next stage.
If results are unsatisfactory, backward iteration may be needed.

FIGURE: The Three-Step/Task Text Mining Process.

Task 1: Establish the Corpus


Definition:
The corpus is a collection of textual documents relevant to the problem domain.
Steps:
1. Collect Textual Data:
o Gather data from multiple sources — documents, emails, web pages, blogs,
XML files, or transcribed voice recordings.
2. Convert to Uniform Format:
o All text must be converted to a consistent digital format (e.g., ASCII or UTF-8
text files).
3. Organize the Corpus:
o Store documents in a structured way (e.g., folders, databases, or linked URLs).
Tools:

Dept of AI&DS, SIET Page 20


BUSINESS ANALYTICS BAD714B

Commercial text mining software can import such data and convert it into flat files for
analysis.
Output:
A well-organized and standardized collection of text documents (the corpus) ready for
processing.

Task 2: Create the Term–Document Matrix (TDM)


Definition:
A Term–Document Matrix (TDM) is a table that represents the relationship between
documents (rows) and terms (columns).
Each cell contains a numerical value (index) showing how often a term appears in a
document.
Document Term 1 Term 2 Term 3 ...
Doc 1 3 0 1 ...
Doc 2 1 5 0 ...

FIGURE: A Simple Term–Document Matrix.

Steps in Creating TDM


1. Text Cleaning and Tokenization
 Split text into individual words or terms (tokens).
 Remove punctuation, numbers, and irrelevant symbols.
2. Stop Word Removal
 Exclude common words (e.g., “a,” “the,” “is”) that don’t contribute meaning.
 These words are called stop words.
3. Stemming and Lemmatization
 Reduce words to their root form (e.g., “modeling,” “modeled” → “model”).
 Ensures that different word forms are treated as the same term.
4. Dictionary Creation
 Define a list of include terms (specific to your study domain).
 Handle synonyms and phrases (e.g., treat “AI” and “Artificial Intelligence” as the
same).
Representing the Indices (Weights)
Once raw term frequencies are computed, the data must be normalized because raw counts
can be misleading.

Dept of AI&DS, SIET Page 21


BUSINESS ANALYTICS BAD714B

Common normalization methods include:


Method Description
Binary 1 if term occurs in document, else 0
Logarithmic scaling to reduce the effect of very
Log Frequency
frequent terms
TF-IDF (Term Frequency–Inverse Increases weight of rare but important terms and
Document Frequency) decreases weight of common terms

Reducing Dimensionality of the TDM


The TDM can become huge and sparse (many zero entries).
Hence, dimensionality reduction is essential for manageable and accurate analysis.
Methods to reduce TDM size:
1. Manual Filtering: Domain experts remove irrelevant or rare terms.
2. Automatic Filtering: Eliminate terms that occur too infrequently.
3. Mathematical Transformation: Apply SVD (Singular Value Decomposition) or
PCA (Principal Component Analysis).
SVD (Singular Value Decomposition):
 Reduces the high-dimensional term space into a few latent dimensions.
 Reveals hidden semantic structures (relationships between words and documents).
 Helps uncover latent semantic meaning — foundation of Latent Semantic Analysis
(LSA).
Output:
A refined and compact TDM, representing documents and their significant terms in
numerical form.

Task 3: Extract the Knowledge


Once the TDM is ready, data mining techniques are applied to discover meaningful
patterns, trends, and relationships.

Main Knowledge Extraction Methods:


[Link] (Text Categorization)
Definition:
Assigns documents to predefined categories (topics, subjects, or sentiment classes).
Examples:
 Email → spam or not spam
 News article → politics, sports, entertainment
 Reviews → positive, negative, neutral
Approaches:
Type Description
Knowledge Engineering Expert defines rules manually.
Machine Learning Algorithms learn from labeled examples (training data).
Popular Algorithms:
Naive Bayes, SVM, Decision Trees, Neural Networks.
Applications:
Automatic indexing, spam filtering, genre detection, web page categorization.

[Link]
Definition:
Groups similar documents together into clusters without predefined labels.

Dept of AI&DS, SIET Page 22


BUSINESS ANALYTICS BAD714B

Goal: Discover natural groupings in data.


Applications:
 Organizing large document repositories (like web pages)
 Enhancing information retrieval and search engines

Benefits:
Type Explanation
Improved Recall Retrieves all relevant documents by considering similar ones.
Improved Precision Returns only the most relevant clusters to the user.
Popular Methods:
 Scatter/Gather Clustering: Dynamically clusters documents for browsing when no
specific query is given.
 Query-Specific Clustering: Hierarchically clusters documents related to a search
query, showing different levels of relevance.

3. Association
Definition:
Discovers relationships or co-occurrences among concepts or terms in documents.
Measures Used:
 Support: % of documents where terms appear together.
 Confidence: Likelihood that one term appears given another.
Example:
If “Software Implementation Failure” frequently appears with “ERP” and “CRM,”
then these terms have a strong association.
Applications:
 Literature analysis (e.g., tracking topics like “bird flu” and related terms such as
“virus,” “vaccine,” “countries affected”).
 Marketing: finding terms that co-occur in customer feedback.

4. Trend Analysis
Definition:
Analyzes how the frequency and relationships of terms evolve over time.
Example:
Tracking how concepts like “AI,” “cloud computing,” or “data privacy” have evolved in
academic journals over years.
Process:
 Compare term distributions across time-based collections.
 Identify emerging or declining topics.
Case Example:
Delen & Crossland (2008) used trend analysis on top information systems journals to
observe how research themes evolved.

Application Case 5.5 Research Literature Survey with Text Mining


Researchers conducting searches and reviews of relevant literature face an increasingly
complex and voluminous task. In extending the body of relevant knowledge, it has always
been important to work hard to gather, organize, analyze, and assimilate existing information
from the literature, particularly from one’s home discipline. With the increasing abundance of
potentially significant research being reported in related fields, and even in what are
traditionally deemed to be nonrelated fields of study, the researcher’s task is ever more
daunting, if a thorough job is desired. In new streams of research, the researcher’s task may

Dept of AI&DS, SIET Page 23


BUSINESS ANALYTICS BAD714B

be even more tedious and complex. Trying to ferret out relevant work that others have
reported may be difficult, at best, and perhaps even near impossible if traditional, largely
manual reviews of published literature are required. Even with a legion of dedicated graduate
students or helpful colleagues, trying to cover all potentially relevant published work is
problematic. Many scholarly conferences take place every year. In addition to extending the
body of knowledge of the current focus of a conference, organizers often desire to offer
additional minitracks and workshops. In many cases, these additional events are intended to
introduce the attendees to significant streams of research in related fields of study and to try
to identify the “next big thing” in terms of research interests and focus. Identifying
reasonable candidate topics for such minitracks and workshops is often subjective rather than
derived objectively from the existing and emerging research. In a recent study, Delen and
Crossland (2008) proposed a method to greatly assist and enhance the efforts of the
researchers by enabling a semiautomated analysis of large volumes of published literature
through the application of text mining. Using standard digital libraries and online publication
search engines, the authors downloaded and collected all the available articles for the three
major journals in the field of management information systems: MIS Quarterly (MISQ),
Information Systems Research (ISR), and the Journal of Management Information Systems
(JMIS). To maintain the same time interval for all three journals (for potential comparative
longitudinal studies), the journal with the most recent starting date for its digital publication
availability was used as the start time for this study (i.e., JMIS articles have been digitally
available since 1994). For each article, they extracted the title, abstract, author list, published
keywords, volume, issue number, and year of publication. They then loaded all the article
data into a simple database file. Also included in the combined data set was a field that
designated the journal type of each article for likely discriminatory analysis. Editorial notes,
research notes, and executive overviews were omitted from the collection. Table 5.2 shows
how the data was presented in a tabular format. In the analysis phase, they chose to use only
the abstract of an article as the source of information extraction. They chose not to include
the keywords listed with the publications for two main reasons: (1) under normal
circumstances, the abstract would already include the listed keywords, and therefore
inclusion of the listed keywords for the analysis would mean repeating the same information
and potentially giving them unmerited weight; and (2) the listed keywords may be terms that
authors would like their article to be associated with (as opposed to what is really contained
in the article), therefore potentially introducing unquantifiable bias to the analysis of the
content. The first exploratory study was to look at the longitudinal perspective of the three
journals (i.e., evolution of research topics over time). In order to conduct a longitudinal study,
they divided the 12-year period (from 1994 to 2005) into four 3-year periods for each of the
three journals. This framework led to 12 text mining experiments with 12 mutually exclusive
data sets. At this point, for each of the 12 data sets they used text mining to extract the most
descriptive terms from these collections of articles represented by their abstracts. The results
were tabulated and examined for time-varying changes in the terms published in these three
journals. As a second exploration, using the complete data set (including all three journals
and all four periods), they conducted a clustering analysis. Clustering is arguably the most
commonly used text mining technique. Clustering was used in this study to identify the
natural groupings of the articles (by putting them into separate clusters) and then to list the
most descriptive terms that characterized those clusters. They used SVD to reduce the
dimensionality of the term-by-document matrix and then an expectation-maximization
algorithm to create the clusters. They conducted several experiments to identify the optimal
number of clusters, which turned out to be nine. After the construction of the nine clusters,
they analyzed the content of those clusters from two perspectives: (1) representation of the
journal type (see Figure 5.8a) and (2) representation of time (Figure 5.8b). The idea was to
explore the potential differences and/or commonalities among the three journals and potential

Dept of AI&DS, SIET Page 24


BUSINESS ANALYTICS BAD714B

changes in the emphasis on those clusters; that is, to answer questions such as “Are there
clusters that represent different research themes specific to a single journal?” and “Is there a
time-varying characterization of those clusters?” They discovered and discussed several
interesting patterns using tabular and graphical representation of their findings (for further
information see Delen & Crossland, 2008).
Table: Tabular Representation of the Fields Included in the Combined Data Set

Questions for Discussion


1. How can text mining be used to ease the insurmountable task of literature review?
2. What are the common outcomes of a text mining project on a specific collection of journal
articles? Can you think of other potential outcomes not mentioned in this case?

Dept of AI&DS, SIET Page 25


BUSINESS ANALYTICS BAD714B

FIGURE: (a) Distribution of the Number of Articles for the Three Journals over the Nine
Clusters; (b) Development of the Nine Clusters over the Years.

Sentiment Analysis
Introduction to Sentiment Analysis
Sentiment Analysis (SA) — also known as Opinion Mining or Subjectivity Analysis — is
a text analytics technique used to identify, extract, and quantify subjective information
such as opinions, emotions, and attitudes expressed in textual data.
In simple terms, sentiment analysis answers the question:
Dept of AI&DS, SIET Page 26
BUSINESS ANALYTICS BAD714B

“What do people feel or think about a certain topic, product, person, or event?”
It transforms unstructured text (like tweets, reviews, posts, and blogs) into structured data
that reflects public sentiment — positive, negative, or neutral.

Background and Importance


Humans naturally depend on others’ opinions — for example:
 Checking reviews before buying a product.
 Reading tweets or blogs before voting or investing.
 Seeking feedback before visiting a restaurant or movie.
Today, people express their opinions widely on:
 Social Media: Twitter, Facebook, Instagram
 Review Platforms: Amazon, IMDb, TripAdvisor
 Forums & Blogs: Reddit, Medium, discussion boards
With the exponential growth of opinion-rich data online, automated tools are required to
analyze massive volumes of text efficiently — this is where sentiment analysis comes in.

Definition and Related Concepts


 Sentiment: Refers to a settled opinion or emotional attitude reflective of one’s
feelings.
(E.g., “This phone is amazing” → positive sentiment)
 Opinion Mining: The process of identifying and classifying opinions expressed in
text.
 Affective Computing: Field focused on recognizing and interpreting human
emotions by computers.
Fields Related to Sentiment Analysis:
 Natural Language Processing (NLP)
 Computational Linguistics
 Machine Learning
 Text Mining

Nature of Sentiments in Text


Sentiments expressed in text can be of two types:
Type Description Example
Explicit Clearly expressed
“The movie was fantastic.”
Sentiment opinion
Implicit Opinion implied “The battery died after one hour.” (implies
Sentiment indirectly negative)
Earlier research focused mainly on explicit sentiments, but modern systems analyze both
for more accuracy.

Polarity in Sentiment Analysis


Sentiment polarity refers to the direction or attitude of opinion.
Type Description Example
Positive Favourable opinion “I love this product.”
Negative Unfavourable opinion “The service was terrible.”
Neutral/Objectivity Factual, no opinion “The phone has a 6-inch screen.”
Sometimes, polarity is measured on a scale (continuum) — e.g.,
very negative → negative → neutral → positive → very positive.

Dept of AI&DS, SIET Page 27


BUSINESS ANALYTICS BAD714B

Applications of Sentiment Analysis


Sentiment Analysis has become a crucial analytical tool across business, politics, and
research domains.
1. Voice of the Customer (VOC)
 Used in Customer Relationship Management (CRM).
 Analyzes customer reviews, emails, feedback, and social media posts.
 Helps companies:
o Detect complaints early (e.g., product issues, bugs).
o Adjust marketing campaigns or advertisements.
o Improve customer satisfaction.
Example:
A movie studio detects negative sentiment about a trailer and changes its marketing strategy
before release.

2. Voice of the Market (VOM)


 Focuses on market-level trends and aggregate opinions.
 Helps with:
o Competitive intelligence
o Product positioning
o Market sentiment tracking
Example: Analyzing social media buzz to compare brand perceptions between Samsung and
Apple.

3. Voice of the Employee (VOE)


 Uses text analytics on employee feedback and surveys.
 Measures employee satisfaction and morale.
 Happy employees → better customer experience → higher productivity.

4. Brand Management
 Monitors online platforms to track brand reputation.
 Detects negative publicity early and manages crises.
 Companies use social listening tools to monitor mentions and opinions.

5. Financial Markets
 Investor sentiment strongly influences stock prices.
 SA helps predict short-term market movements based on:
o News headlines
o Social media buzz
o Discussion forums
 Example: Detecting market panic or optimism before trading decisions.

6. Politics
 Analyzes public opinions during elections or debates.
 Predicts voter preferences, popularity of candidates, and reaction to policies.
 Used successfully in U.S. presidential campaigns (2008, 2012).

7. Government Intelligence
 Monitors public reactions to policies, proposals, or events.
 Detects hostile or negative sentiment spikes for security or regulatory alerts.
 Useful for agencies like Homeland Security.

Dept of AI&DS, SIET Page 28


BUSINESS ANALYTICS BAD714B

8. Other Interesting Areas


 E-commerce: Personalized recommendations based on sentiment.
 Email Filtering: Detect negative/urgent messages automatically.
 Ad Placement: Display relevant ads based on user sentiment.
 Academic Research: Identify supportive or critical citations.

Sentiment Analysis Process (Step-by-Step)


There’s no single standardized process, but a four-step iterative workflow is widely
followed.
Step 1: Sentiment Detection
 Distinguish between facts and opinions (Objective vs. Subjective sentences).
 Compute Objectivity–Subjectivity (O–S) polarity between 0 and 1.
o 1 → Objective (fact)
o 0 → Subjective (opinion)
 Typically, adjectives are key indicators:
o “Wonderful performance” → subjective → positive sentiment.
Step 2: Polarity Classification (N–P Polarity)
 Classify opinionated text as Positive or Negative (binary classification).
 Sometimes includes intensity/strength (e.g., mild, moderate, strong).
 Works at different levels of granularity:
o Word level
o Sentence level
o Document level
Challenges:
 Mixed sentiments in one document.
 Implicit emotions (e.g., sarcasm or irony).
Common Algorithms:
Naive Bayes, SVM, Neural Networks, Decision Trees.
Step 3: Target Identification
 Identify who or what the sentiment is directed toward.
o e.g., “I love Apple’s design but hate their prices.” → two targets.
 Comparative sentences are handled using comparative words:
o “This laptop is better than that desktop.”
Goal: Match the sentiment to the correct entity or aspect.
Step 4: Collection and Aggregation
 Combine all detected sentiments into a single summary measure.
 Aggregation methods:
o Simple averaging or summing of sentiment scores.
o Semantic aggregation (context-aware methods).
Example:
If a product review has 5 positive and 2 negative opinions → overall sentiment = Positive.

Dept of AI&DS, SIET Page 29


BUSINESS ANALYTICS BAD714B

FIGURE: A Multistep Process to Sentiment Analysis.


Methods for Polarity Identification:
Polarity can be determined using two main approaches:
A. Lexicon-Based Approach
Uses predefined dictionaries (lexicons) that list words and their sentiment scores.
Common Lexicons:
1. WordNet:
o A large lexical database grouping words into synsets (synonym sets).
o Forms the foundation for extensions like SentiWordNet and WordNet-
Affect.
2. SentiWordNet:
o Assigns positivity, negativity, and objectivity scores (0.0 to 1.0) to each
synset.
o Example:
 “Good” → Positive 0.8, Negative 0.0, Objective 0.2.
3. WordNet-Affect:
o Adds emotion-related categories like anger, joy, or fear to WordNet.
4. Custom Domain Lexicons:
o Created for specific industries (e.g., finance, medicine).
Advantages:
 Simple and interpretable.
 Works well for rule-based systems.
Disadvantages:

Dept of AI&DS, SIET Page 30


BUSINESS ANALYTICS BAD714B

 Ignores context (e.g., sarcasm).


 Limited vocabulary for domain-specific terms.

FIGURE: A Graphical Representation of the P–N Polarity and S–O Polarity


Relationship

B. Machine Learning Approach


Uses labeled datasets to train predictive models that learn patterns of sentiment
automatically.
Data Sources for Training:
 Product review sites: Amazon, IMDb, eBay, RottenTomatoes.
 Benchmark datasets:
o Text REtrieval Conference (TREC)
o Cross-Language Evaluation Forum (CLEF)
o NII Test Collection for IR Systems
Algorithms Used:
 Naive Bayes
 Support Vector Machines (SVM)
 Artificial Neural Networks
 k-Nearest Neighbors (kNN)
 Decision Trees
 Expectation–Maximization Clustering
Advantages:
 Automatically adapts to new data.
 Domain-specific accuracy.
Disadvantages:
 Requires large labeled datasets.
 Can be computationally expensive.

Semantic Orientation of Words, Phrases, and Documents


1. Word-Level
Each word is assigned a polarity score (positive, negative, objective).
2. Phrase/Sentence-Level
Aggregate word scores to find sentiment of a phrase or sentence.
Example:
“The camera quality is amazing but battery life is poor.”
→ Positive for camera, negative for battery.
3. Document-Level
Average or aggregate sentiment across the document to obtain overall orientation.
Useful for:

Dept of AI&DS, SIET Page 31


BUSINESS ANALYTICS BAD714B

 Short to medium-length documents (e.g., reviews, posts).


 Not suitable for very long texts with multiple topics.

Visualization and Business Use


Modern sentiment analysis tools integrate results into dashboards that:
 Display real-time customer emotions.
 Track brand reputation trends.
 Enable data-driven marketing and policy decisions.

Application Case 5.6: Creating a Unique Digital Experience to Capture the Moments
That Matter at Wimbledon

Known to millions of fans simply as “Wimbledon,” The Championships is the oldest of


tennis’ four Grand Slams, and one of the world’s highest-profile sporting events. Organized
by the All England Lawn Tennis Club (AELTC) it has been a global sporting and cultural
institution since 1877.
The Champion of Championships
The organizers of The Championships, Wimbledon, the AELTC, have a simple objective:
every year, they want to host the best tennis championships in the world—in every way, and
by every metric. The motivation behind this commitment is not simply pride; it also has a
commercial basis. Wimbledon’s brand is built on its premier status: this is what attracts both
fans and partners. The world’s best media organizations and greatest corporations— IBM
included—want to be associated with Wimbledon precisely because of its reputation for
excellence. For this reason, maintaining the prestige of The Championships is one of the
AELTC’s top priorities, but there are only two ways that the organization can directly control
how The Championships are perceived by the rest of the world. The first, and most important,
is to provide an outstanding experience for the players, journalists, and spectators who are
lucky enough to visit and watch the tennis courtside. The AELTC has vast experience in this
area. Since 1877 it has delivered two weeks of memorable, exciting competition in an idyllic
setting: tennis in an English country garden. The second is The Championships’ online
presence, which is delivered via the [Link] Website, mobile apps, and social media

Dept of AI&DS, SIET Page 32


BUSINESS ANALYTICS BAD714B

channels. The constant evolution of these digital platforms is the result of a 26-year
partnership between the AELTC and IBM. Mick Desmond, Commercial and Media Director
at the AELTC, explains: “When you watch Wimbledon on TV, you are seeing it through the
broadcaster’s lens. We do everything we can to help our media partners put on the best
possible show, but at the end of the day, their broadcast is their presentation of The
Championships. “Digital is different: it’s our platform, where we can speak directly to our
fans—so it’s vital that we give them the best possible experience. No sporting event or media
channel has the right to demand a viewer’s attention, so if we want to strengthen our brand,
we need people to see our digital experience as the number-one place to follow The
Championships online.” To that end, the AELTC set a target of attracting 70 million visits, 20
million unique devices, and 8 million social followers during the two weeks of The
Championships 2015. It was up to IBM and AELTC to find a way to deliver.
Delivering a Unique Digital Experience
IBM and the AELTC embarked on a complete redesign of the digital platform, using their
intimate knowledge of The Championships’ audience to develop an experience tailor-made to
attract and retain tennis fans from across the globe. “We recognized that while mobile is
increasingly important, 80% of our visitors are using desktop computers to access our
website,” says Alexandra Willis, Head of Digital and Content at the AELTC. “Our challenge
for 2015 was how to update our digital properties to adapt to a mobilefirst world, while still
offering the best possible desktop experience. We wanted our new site to take maximum
advantage of that large screensize and give desktop users the richest possible experience in
terms of high-definition visuals and video content—while also reacting and adapting
seamlessly to smaller tablet or mobile formats. “Second, we placed a major emphasis on
putting content in context—integrating articles with relevant photos, videos, stats and
snippets of information, and simplifying the navigation so that users could move seamlessly
to the content that interests them most.” On the mobile side, the team recognized that the
wider availability of high-bandwidth 4G connections meant that the mobile Website would
become more popular than ever—and ensured that it would offer easy access to all rich media
content. At the same time, The Championships’ mobile apps were enhanced with real-time
notifications of match scores and events—and could even greet visitors as they passed
through stations on the way to the grounds. The team also built a special set of Websites for
the most important tennis fans of all: the players themselves. Using IBM® Bluemix®
technology, it built a secure Web application that provided players with a personalized view
of their court bookings, transport, and on-court times, as well as helping them review their
performance with access to stats on every match they played.
Turning Data into Insight—and Insight into Narrative
To supply its digital platforms with the most compelling possible content, the team took
advantage of a unique advantage: its access to real-time, shotby-shot data on every match
played during The Championships. Over the course of the Wimbledon fortnight, 48 courtside
experts capture approximately 3.4 million datapoints, tracking the type of shot, the strategies,
and the outcome of each and every point. This data is collected and analyzed in real time to
produce statistics for TV commentators and journalists—and also for the digital platform’s
own editorial team. “This year IBM gave us an advantage that we had never had before—
using data streaming technology to provide our editorial team with real-time insight into
significant milestones and breaking news,” says Alexandra Willis. “The system automatically
watched the streams of data coming in from all 19 courts, and whenever something
significant happened—such as Sam Groth hitting the second-fastest serve in Championships,
history—it let us know instantly. Within seconds, we were able to bring that news to our
digital audience and share it on social media to drive even more traffic to our site. “The
ability to capture the moments that matter and uncover the compelling narratives within the
data, faster than anyone else, was key. If you wanted to experience the emotions of The

Dept of AI&DS, SIET Page 33


BUSINESS ANALYTICS BAD714B

Championships live, the next best thing to being there in person was to follow the action on
[Link].”
Harnessing the Power of Natural Language
Another new capability trialed this year was the use of IBM’s NLP technologies to help mine
the AELTC’s huge library of tennis history for interesting contextual information. The team
trained IBM Watson™ Engagement Advisor to digest this rich unstructured data set and use
it to answer queries from the press desk. The same NLP front-end was also connected to a
comprehensive structured database of match statistics, dating back to the first Championships
in 1877—providing a one-stop shop for both basic questions and more complex inquiries.
“The Watson trial showed a huge amount of potential. Next year, as part of our annual
innovation planning process, we will look at how we can use it more widely—ultimately in
pursuit of giving fans more access to this incredibly rich source of tennis knowledge,” says
Mick Desmond.
Taking to the Cloud
The whole digital environment was hosted by IBM in its Hybrid Cloud. IBM used
sophisticated modeling techniques to predict peaks in demand based on the schedule, the
popularity of each player, the time of day, and many other factors—enabling it to
dynamically allocate cloud resources appropriately to each piece of digital content and ensure
a seamless experience for millions of visitors around the world. In addition to the powerful
private cloud platform that has supported The Championships for several years, IBM also
used a separate SoftLayer® cloud to host the Wimbledon Social Command Centre and also
provide additional incremental capacity to supplement the main cloud environment during
times of peak demand. The elasticity of the cloud environment is key, as The Championships’
digital platforms need to be able to scale efficiently by a factor of more than 100 within a
matter of days as the interest builds ahead of the first match on Centre Court.
Keeping Wimbledon Safe and Secure
Online security is a key concern nowadays for all organizations. For major sporting events in
particular, brand reputation is everything—and while the world is watching, it is particularly
important to avoid becoming a high-profile victim of cyber-crime. For these reasons, security
has a vital role to play in IBM’s partnership with the AELTC. Over the first five months of
2015, IBM security systems detected a 94% increase in security events on the
[Link] infrastructure, compared to the same period in 2014. As security threats—
and in particular distributed denial of service (DDoS) attacks—become ever more prevalent,
IBM continually increases its focus on providing industry-leading levels of security for the
AELTC’s whole digital platform. A full suite of IBM security products, including IBM
QRadar® SIEM and IBM Preventia Intrusion Prevention, enabled this year’s Championships
to run smoothly and securely and the digital platform to deliver a high-quality user
experience at all times.
Capturing Hearts and Minds
The success of the new digital platform for 2015— supported by IBM cloud, analytics,
mobile, social, and security technologies—was immediate and complete. Targets for total
visits and unique visitors were not only met, but exceeded. Achieving 71 million visits and
542 million page views from 21.1 million unique devices demonstrates the platform’s success
in attracting a larger audience than ever before and keeping those viewers engaged
throughout The Championships. “Overall, we had 13% more visits from 23% more devices
than in 2014, and the growth in the use of [Link] on mobile was even more
impressive,” says Alexandra Willis. “We saw 125% growth in unique devices on mobile,
98% growth in total visits, and 79% growth in total page views.” Mick Desmond concludes:
“The results show that in 2015, we won the battle for fans’ hearts and minds. People may
have favorite newspapers and sports websites that they visit for 50 weeks of the year—but for
two weeks, they came to us instead. “That’s a testament to the sheer quality of the experience

Dept of AI&DS, SIET Page 34


BUSINESS ANALYTICS BAD714B

we can provide—harnessing our unique advantages to bring them closer to the action than
any other media channel. The ability to capture and communicate relevant content in real
time helped our fans experience The Championships more vividly than ever before.”
Questions for Discussion
1. How did Wimbledon use analytics capabilities to enhance viewers’ experience?
2. What were the challenges, the proposed solution, and the obtained results?

Topic Modeling
Definition:
Topic modeling is an advanced text mining technique used to automatically discover
hidden thematic structures (topics) from a large collection of unstructured text documents.
It identifies groups of words that frequently occur together and represent a particular concept
or theme within the corpus.
Purpose of Topic Modeling
 To organize and summarize large collections of textual data.
 To uncover abstract “topics” discussed across multiple documents.
 To help in document classification, summarization, and recommendation systems.
 To support decision-making by identifying key themes, trends, or patterns in text data.

How It Works
Topic modeling assumes that:
 Each document is composed of multiple topics.
 Each topic is characterized by a distribution of words.
Example:
Document: “The movie had great visual effects but poor acting.”
Topics identified:
• Entertainment (words: movie, acting, performance)
• Technology (words: visual, effects, graphics)

Main Algorithms for Topic Modeling


A. Latent Semantic Analysis (LSA)
 Based on Singular Value Decomposition (SVD).
 Reduces dimensionality of the term–document matrix.
 Reveals patterns in word usage and relationships between terms and documents.
 Limitation: LSA captures linear relationships only and is less effective for
probabilistic modeling.
B. Probabilistic Latent Semantic Analysis (PLSA)
 Improves on LSA by introducing probabilistic models.
 Each document is modeled as a mixture of topics with certain probabilities.
 Each topic is represented as a distribution over words.
 Formula:

 where z represents topics.


 Limitation: May overfit due to the lack of a proper prior distribution.

Dept of AI&DS, SIET Page 35


BUSINESS ANALYTICS BAD714B

C. Latent Dirichlet Allocation (LDA)


 Most popular and widely used topic modeling algorithm.
 Introduced by Blei, Ng, and Jordan (2003).
 A generative probabilistic model that assumes:
o Each document is a mixture of multiple topics.
o Each topic is a mixture of words.
o Both are governed by Dirichlet distributions (priors).
LDA Steps:
1. Randomly assign topics to words in documents.
2. Iteratively update topic assignments to improve fit based on word co-occurrences.
3. Estimate:
o P(topic ∣ document)
o P(word ∣ topic)
4. Output: A set of topics with their key terms and topic proportions per document.
Advantages:
 Handles large corpora efficiently.
 Produces interpretable and coherent topics.
 Can be extended (e.g., Dynamic LDA, Correlated Topic Models).
Applications of LDA:
 Identifying research trends in academic papers.
 Social media text analysis.
 Customer feedback and review mining.
 News topic classification.

Steps in Topic Modeling Process


1. Data Collection
o Gather text documents (e.g., news, reviews, articles, reports).
2. Text Preprocessing
o Tokenization
o Stop-word removal
o Lemmatization/Stemming
o Creating the term–document matrix (TDM) or document–term matrix
(DTM).
3. Model Building
o Apply topic modeling algorithms (e.g., LDA, NMF, PLSA).
o Specify the number of topics (K) to discover.
4. Model Evaluation
o Evaluate topic coherence (semantic similarity between top words in a topic).
o Adjust number of topics to optimize interpretability.
5. Visualization
o Use word clouds, topic-term distributions, and heatmaps.
o Tools like pyLDAvis, Gensim, or Tableau can visualize topic relationships.

Applications of Topic Modeling

Dept of AI&DS, SIET Page 36


BUSINESS ANALYTICS BAD714B

 Business Intelligence: Extract customer concerns or product themes from reviews.


 Healthcare: Identify emerging disease research topics.
 Education: Analyze student feedback and academic literature.
 Social Media: Monitor public sentiment and trending issues.
 Law Enforcement: Analyze reports or online content for criminal patterns.

Tools and Libraries for Topic Modeling


 Python: Gensim, scikit-learn, spaCy, pyLDAvis.
 R: topicmodels, LDAvis.
 Big Data Platforms: Apache Spark (MLlib), RapidMiner, KNIME.

Limitations of Topic Modeling


 Choosing the right number of topics (K) can be subjective.
 Topics may not always be interpretable.
 Sensitive to preprocessing (especially stop-word handling).
 Assumes “bag-of-words,” ignoring word order and context (though extensions like
BERT-based topic modeling now address this).

Web Mining Overview


Introduction
The Internet has transformed how businesses operate, interact with customers, and compete
in a global marketplace.
Today, the Web is not only a communication medium but also a massive data source
containing structured, semi-structured, and unstructured information.
Modern organizations rely heavily on Web mining to analyze this information and gain
competitive advantages through insights into:
 Customer preferences
 Market trends
 Competitor activities
 Online behavior and sentiments

Why Web Mining Is Important


 Businesses must maintain a strong online presence to reach customers effectively.
 Consumers share feedback, opinions, and experiences through websites, forums,
and social media.
 Data-driven decision making requires extracting useful patterns from large volumes
of Web data.

Definition of Web Mining


Web Mining (or Web Data Mining) is the process of discovering intrinsic patterns and
useful information from Web-related data — including text, hyperlinks, and usage logs.
Simply put:
Web mining = Data mining techniques applied to Web data.
Introduced by Etzioni (1996), Web mining combines techniques from:

Dept of AI&DS, SIET Page 37


BUSINESS ANALYTICS BAD714B

 Data mining
 Text mining
 Information retrieval
 Machine learning
 Natural language processing (NLP)

Difference Between Web Mining and Web Analytics


Aspect Web Mining Web Analytics
Broader scope – includes content, structure, Focused mainly on website usage and
Focus
and usage data visitor metrics
Discover hidden patterns, trends, and Measure and report “what happened”
Goal
relationships on a website
Predictive & prescriptive (insight
Nature Descriptive (metrics and summaries)
discovery)
Discovering patterns in customer Tracking page views, bounce rates,
Example
navigation session times
In short: Web analytics is part of Web mining.

Challenges in Web Mining


Mining the Web is difficult due to its massive, complex, and dynamic nature:
1. Too Big: The Web’s size makes it impossible to collect or store all data in a single
database.
2. Too Complex: Web pages vary in format (HTML, XML, multimedia) and lack a unified
structure.
3. Too Dynamic: Content changes rapidly — blogs, news, social media, and e-commerce
prices are updated constantly.
4. Too Diverse: Web users come from different domains, interests, and expertise levels.
5. Too Noisy: Only a small portion of the Web is useful to a specific user; most of it is
irrelevant or redundant.
Taxonomy of Web Mining
Web mining can be categorized into three main areas:
Category Data Source Purpose Typical Techniques / Tools
1. Web Web page content Extract information from Text mining, NLP, Web
Content (text, images, videos, the content of Web crawlers, classification,
Mining etc.) documents summarization
2. Web Analyze structure of the
Hyperlinks (inter- Graph theory, link analysis
Structure Web to identify
page links) (HITS, PageRank)
Mining authorities and hubs
Analyze user behavior Data mining, clustering,
3. Web Usage Web server logs,
and interactions with association rule learning,
Mining clickstreams
websites session analysis

Dept of AI&DS, SIET Page 38


BUSINESS ANALYTICS BAD714B

FIGURE: A Simple Taxonomy of Web Mining.

1. Web Content Mining


Definition
Web content mining is the process of extracting useful information from the content of
Web pages — including text, images, audio, and video.
Data Sources
 Static web pages (HTML, XML)
 Blogs, news articles
 Product reviews
 Multimedia content (images, videos)
 Social media posts
Techniques Used
 Web crawlers/spiders: Automated programs that scan and collect data from web
pages.
 Text mining: Keyword extraction, sentiment analysis, and summarization.
 Information extraction: Identifying entities and relationships.
 Classification and clustering: Grouping similar content or topics.
Applications
 Competitive intelligence (tracking competitor websites)
 Automated data collection (e.g., movie or product data)
 Opinion mining and sentiment analysis
 Content recommendation systems
Example

Dept of AI&DS, SIET Page 39


BUSINESS ANALYTICS BAD714B

Researchers use web crawlers to automatically collect data about thousands of movies —
including box office revenue, genre, cast, and ratings — to predict financial success using
data mining models.

2. Web Structure Mining


Definition
Web structure mining focuses on analyzing the link structures within and between web
pages to understand the relationship among them.
Key Concepts
 Hyperlinks: Connections between web pages (a form of implicit human judgment).
 Authority: A page that many others link to — considered a credible or influential
source.
 Hub: A page that links to many authoritative pages on a particular topic.
A good hub points to many good authorities; a good authority is linked by many good hubs.
Purpose
 Identify important or authoritative web pages.
 Enhance search engine ranking.
 Discover communities and clusters of related websites.
Algorithms
1. HITS (Hyperlink-Induced Topic Search) – by Jon Kleinberg (1999):
o Calculates two scores for each page:
 Hub score (quality of outgoing links)
 Authority score (quality of incoming links)
o Works recursively until scores stabilize.
2. PageRank (used by Google):
o Evaluates the importance of a page based on the number and quality of
incoming links.
Applications
 Search engine optimization (SEO)
 Identifying web communities or influence networks
 Ranking web pages based on importance or topic relevance

3. Web Usage Mining


Definition
Web usage mining is the analysis of web logs and user activity data to understand and
predict user behavior.
Data Sources
 Web server access logs
 Browser cookies
 Clickstream data
 Session histories
Applications
 Personalization and recommendation systems
 Website optimization

Dept of AI&DS, SIET Page 40


BUSINESS ANALYTICS BAD714B

 Fraud detection
 Predicting customer churn or conversion
Typical Techniques
 Clustering: Group users with similar browsing patterns.
 Association rules: Discover frequently visited page sequences.
 Sequential pattern mining: Identify navigation paths.

Process of Web Mining


Step Description
1. Data Collection Using crawlers, log files, APIs, etc., to gather data.
2. Preprocessing Cleaning, formatting, and transforming raw Web data.
3. Pattern
Applying data mining, machine learning, and NLP algorithms.
Discovery
4. Pattern Analysis Evaluating and interpreting discovered knowledge.
Using insights for decision-making, personalization, or business
5. Application
intelligence.

Applications of Web Mining


1. E-commerce: Product recommendations, cross-selling, churn prediction.
2. Marketing: Customer segmentation, sentiment tracking, brand monitoring.
3. Finance: Fraud detection and market intelligence.
4. Search Engines: Ranking, categorization, and topic discovery.
5. Social Media Analytics: Influence detection and opinion mining.
6. Healthcare: Disease trend monitoring from online sources.
7. Education: Analyzing online learning behavior and content usage.

Advantages
 Extracts actionable knowledge from the vast Web.
 Enables real-time business decision making.
 Enhances personalization and customer experience.
 Improves competitiveness through data-driven insights.

Limitations
 Privacy and ethical concerns (tracking user data).
 Dynamic nature of Web content leads to data inconsistency.
 High computational cost and storage requirements.
 Data quality issues (spam, redundancy, or missing metadata).

Search Engine
A search engine is a software system that searches information stored on the Web and
returns results based on the user’s query.
Examples: Google, Bing, Yahoo, DuckDuckGo.

Dept of AI&DS, SIET Page 41


BUSINESS ANALYTICS BAD714B

When users type keywords or a sentence, the search engine returns Web pages that are most
relevant to those keywords.
Goals of a Search Engine
Goal Meaning
Effectiveness (Quality) Return the most relevant and useful search results.
Efficiency (Speed) Return results quickly.
These two goals are balanced to provide the best user experience.

Why Are Search Engines Important?


 The Internet is huge and complex.
 Without search engines, it would be nearly impossible to find useful information.
 They are used for:
o Researching products and services
o Finding locations, people, events
o Shopping and comparing prices
o Learning, entertainment, education, etc.
Search engines are the gateway to the digital world.

Anatomy / Structure of a Search Engine


A search engine operates in two main cycles:
Cycle Purpose
Development Cycle Builds and updates the search engine’s page database.
Response Cycle Answers user queries in real time.

A. Development Cycle
It includes:
1. Web Crawler (Spider)
2. Document Indexer
i. Web Crawler
 Automatically browses and downloads Web pages.
 Starts from a list of URLs (called seeds).
 Finds new pages by following hyperlinks.
 Stores these pages for processing.

FIGURE: Structure of a Typical Internet Search Engine.

Dept of AI&DS, SIET Page 42


BUSINESS ANALYTICS BAD714B

ii. Document Indexer


Processes fetched pages and creates searchable index.
Steps:
1. Convert documents into standard format.
2. Parse text → break into tokens (words).
3. Remove stop words (e.g., “is, the, a, an…”).
4. Stemming → Convert words to base form (e.g., “playing” → “play”).
5. Create Term-by-Document matrix (often using TF/IDF weighting).
This structured indexing enables fast retrieval later.

B. Response Cycle
It includes:
1. Query Analyzer
2. Document Matcher/Ranker
i. Query Analyzer
 Takes user’s search text.
 Converts it to same structure used in indexing.
 Performs:
o Tokenization
o Stop-word removal
o Stemming
o Spell correction / synonym checks
ii. Document Matcher and Ranker
 Matches processed query against indexed database.
 Ranks documents by relevance.

Page Ranking
Originally, search engines only matched keywords — quality was poor.
Google introduced PageRank Algorithm (1997):
 A page is important if many other important pages link to it.
 Similar to academic citation: highly cited pages are more influential.
This ranking, combined with content relevance, improves search result quality.

Search Engine Optimization (SEO)


SEO = Improving a website so that search engines rank it higher.
Higher ranking → More visibility → More visitors → More sales.
Two Types of SEO
Type Description Example Techniques
White-Hat Ethical & search-engine-approved Quality content, keyword optimization,
SEO methods clean HTML
Black-Hat Unethical shortcuts that manipulate
Hidden text, keyword stuffing, cloaking
SEO ranking
Black-hat SEO may lead to search engine penalties or banning.

Dept of AI&DS, SIET Page 43


BUSINESS ANALYTICS BAD714B

SEO Techniques
Technique Purpose
Cross-linking pages internally Strengthens important pages
Use relevant keywords in content Helps match user queries
Updating content regularly Encourages frequent crawling
Adding keywords in title and meta descriptions Helps search engine relevance
Increasing backlinks from other sites Improves authority and PageRank

Challenges in Search Engines


 Web is huge and growing rapidly.
 Web pages lack uniform structure.
 Web content changes frequently.
 Search engines must constantly adapt algorithms.
Google alone makes ~500+ changes per year.

Application Case 5.7: Understanding Why Customers Abandon Shopping Carts Results
in a $10 Million Sales Increase
[Link], the leading Internet shopping mall in Korea with 13 million customers, has
developed an integrated Web traffic analysis system using SAS for Customer Experience
Analytics. As a result, Lotte. com has been able to improve the online experience for its
customers, as well as generate better returns from its marketing campaigns. Now, Lotte. com
executives can confirm results anywhere, anytime, as well as make immediate changes. With
almost one million Web site visitors each day, [Link] needed to know how many visitors
were making purchases and which channels were bringing the most valuable traffic. After
reviewing many diverse solutions and approaches, [Link] introduced its integrated Web
traffic analysis system using the SAS for Customer Experience Analytics solution. This is the
first online behavioral analysis system applied in Korea. With this system, [Link] can
accurately measure and analyze Web site visitor numbers, page view status of site visitors
and purchasers, the popularity of each product category and product, clicking preferences for
each page, the effectiveness of campaigns, and much more. This information enables
[Link] to better understand customers and their behavior online, and conduct
sophisticated, costeffective targeted marketing. Commenting on the system, Assistant
General Manager Jung Hyo-hoon of the Marketing Planning Team for [Link] said, “As a
result of introducing the SAS system of analysis, many ‘new truths’ were uncovered around
customer behavior, and some of them were ‘inconvenient truths.’” He added, “Some site-
planning activities that had been undertaken with the expectation of certain results actually
had a low reaction from customers, and the site planners had a difficult time recognizing
these results.”
Benefits
Introducing the SAS for Customer Experience Analytics solution fully transformed the
[Link] Web site. As a result, [Link] has been able to improve the online experience

Dept of AI&DS, SIET Page 44


BUSINESS ANALYTICS BAD714B

for its customers as well as generate better returns from its marketing campaigns. Since
implementing SAS for Customer Experience Analytics, [Link] has seen many benefits.
A Jump in Customer Loyalty
A large amount of sophisticated activity information can be collected under a visitor
environment, including quality of traffic. Deputy Assistant General Manager Jung said that
“by analyzing actual valid traffic and looking only at one to two pages, we can carry out
campaigns to heighten the level of loyalty, and determine a certain range of effect,
accordingly.” He added, “In addition, it is possible to classify and confirm the order rate for
each channel and see which channels have the most visitors.”
Optimized Marketing Efficiency Analysis
Rather than just analyzing visitor numbers only, the system is capable of analyzing the
conversion rate (shopping cart, immediate purchase, wish list, purchase completion)
compared to actual visitors for each campaign type (affiliation or e-mail, banner, keywords,
and others), so detailed analysis of channel effectiveness is possible. In addition, it can
confirm the most popular search words used by visitors for each campaign type, location, and
purchased products. The page overlay function can measure the number of clicks and number
of visitors for each item in a page to measure the value for each location in a page. This
capability enables [Link] to promptly replace or renew low-traffic items.
Enhanced Customer Satisfaction and Customer Experience Lead to Higher Sales
[Link] built a customer behavior analysis database that measures each visitor, what pages
are visited, how visitors navigate the site, and what activities are undertaken to enable diverse
analysis and improve site efficiency. In addition, the database captures customer
demographic information, shopping cart size and conversion rate, number of orders, and
number of attempts. By analyzing which stage of the ordering process deters the most
customers and fixing those stages, conversion rates can be increased. Previously, analysis
was done only on placed orders. By analyzing the movement pattern of visitors before
ordering and at the point where breakaway occurs, customer behavior can be forecast, and
sophisticated marketing activities can be undertaken. Through a pattern analysis of visitors,
purchases can be more effectively influenced and customer demand can be reflected in real
time to ensure quicker responses. Customer satisfaction has also improved as Lotte. com has
better insight into each customer’s behaviors, needs, and interests. Evaluating the system,
Jung commented, “By finding out how each customer group moves on the basis of the data, it
is possible to determine customer service improvements and target marketing subjects, and
this has aided the success of a number of campaigns.” However, the most significant benefit
of the system is gaining insight into individual customers and various customer groups. By
understanding when customers will make purchases and the manner in which they navigate
throughout the Web page, targeted channel marketing and better customer experience can
now be achieved. Plus, when SAS for Customer Experience Analytics was implemented by
[Link]’s largest overseas distributor, it resulted in a first-year sales increase of 8 million
euros (US$10 million) by identifying the causes of shopping cart abandonment.
Questions for Discussion
1. How did [Link] use analytics to improve sales?
2. What were the challenges, the proposed solution, and the obtained results?

Dept of AI&DS, SIET Page 45


BUSINESS ANALYTICS BAD714B

3. Do you think e-commerce companies are in better position to leverage benefits of


analytics? Why? How?

Web Usage Mining (Web Analytics)


Web usage mining—often referred to as Web analytics—is the process of extracting useful
patterns and knowledge from data generated when users interact with websites.
This includes data from:
 Page visits
 Clicks
 Navigation paths
 Search queries
 Online purchases
 Download activities
The goal is to understand user behavior and use this knowledge to improve website
performance, marketing effectiveness, and customer experience.

What Is Web Usage Mining?


Web usage mining focuses on the log data stored by Web servers, browsers, proxy servers,
and online transactions.
Example:
A visitor searches “Hotels in Goa” and then searches “Flights to Goa”.
This pattern can help travel websites place flight ads alongside hotel search pages.
This type of behavioral pattern analysis is called Clickstream Analysis.

Why Is Web Usage Mining Important?


It helps organizations:
 Understand what users do on their websites.
 Identify popular pages and navigation patterns.
 Improve website design and content.
 Target advertisements and promotions more effectively.
 Increase sales, engagement, and customer satisfaction.

FIGURE: Extraction of Knowledge from Web Usage Data.

Dept of AI&DS, SIET Page 46


BUSINESS ANALYTICS BAD714B

Web Analytics Technologies


Web analytics refers to measuring, collecting, analyzing, and interpreting Web usage
data.
Two Main Types of Web Analytics
Type Meaning Example
Off-site Web Analysis of a brand's presence Tracking brand mentions on social
Analytics outside its own website. media, blogs, forums.
On-site Web Analysis of how users behave on Tracking page views, purchases,
Analytics the website itself. clicks, navigation paths.

Two Technical Approaches for On-site Data Collection


Method How It Works Advantage
Server Log File Web server logs store every request No need to change site
Analysis made. pages.
Page Tagging JavaScript code embedded in webpages More accurate tracking of
(JavaScript Tags) sends data to analytics servers. user actions & events.
Common tools: Google Analytics, Adobe Analytics, Matomo, Yahoo Analytics, Microsoft
Clarity.

Key Metrics in Web Analytics


Web analytics metrics are grouped into four main categories:
A. Website Usability Metrics → How users interact with the site
Metric Meaning Importance
Page Views Number of pages viewed Low page views → poor content or structure
Time on Site How long visitors stay More time → more engagement
Downloads Number of resources downloaded Shows interest level
Click Map Shows where users click Identifies important vs. ignored elements
Click Paths Navigation routes taken Helps improve user journey based on goals

B. Traffic Source Metrics → Where visitors come from


Source Description
Referral Websites Other websites linking to your site
Search Engines Visitors arriving via Google/Bing searches
Direct Traffic User types URL directly or uses bookmarks
Offline Campaigns Flyers, TV ads directing to special URLs
Online Campaigns Email, social media, banner ads, Google Ads

C. Visitor Profile Metrics → Who the visitors are


Metric Insight Gained
Keywords Used Understand what users are searching for

Dept of AI&DS, SIET Page 47


BUSINESS ANALYTICS BAD714B

Metric Insight Gained


Content Grouping Which type of content is most popular
Geography Cities/countries generating most traffic
Time of Day When visitors are most active
Landing Page Which pages users enter first

D. Conversion Statistics → Whether business goals are achieved


A "conversion" = User takes desired action (e.g., purchase, form submission, registration,
etc.)
Metric Meaning
New Visitors Measures brand awareness
Returning Visitors Measures loyalty and engagement
Leads Number of inquiries or sign-ups
Sales Conversions Number of purchases or transactions
Abandonment/Exit Rates Where users leave the site → Indicates problem points

How Web Usage Mining Helps Businesses


 Improves website layout to reduce user drop-off.
 Optimizes marketing campaigns using real-time performance data.
 Targets customers with relevant promotions based on browsing behavior.
 Boosts conversion rates and overall revenue.
 Strengthens customer satisfaction and loyalty.

FIGURE: A Sample Web Analytics Dashboard.

Dept of AI&DS, SIET Page 48


BUSINESS ANALYTICS BAD714B

Social Analytics
Social Analytics refers to the monitoring, analyzing, measuring, and interpreting digital
interactions among people on social platforms. Its purpose is to extract business insights
from social behavior and communication patterns.
It focuses on:
 What people say (text, comments, posts, reviews)
 How people interact (likes, shares, follows, influence)
 How groups form and evolve online
According to Gartner, social analytics includes:
 Measuring conversations
 Tracking relationships
 Understanding community influence
 Predicting behaviors and trends

Two Major Branches of Social Analytics


Branch Focus Area Techniques Used Goal
1. Social Structure of
Graph theory, network Identify influencers,
Network relationships between
metrics community patterns
Analysis (SNA) people/groups
Content and interactions Text mining, sentiment Understand customer
2. Social Media
on social media analysis, engagement attitudes and improve
Analytics
platforms metrics marketing

Social Network Analysis (SNA)


What Is a Social Network?
A social network is a structure made of:
 Nodes → Individuals, groups, or organizations
 Edges (Ties) → Relationships or interactions (friendship, communication,
collaboration)
SNA studies how these networks form, spread information, and influence behavior.
Why Is SNA Important?
Organizations use SNA to:
 Identify key influencers (opinion leaders)
 Understand information flow
 Detect clusters or communities
 Study behavioral patterns within groups
Common Types of Social Networks
Type Example Business Use
Communication Email/phone interactions in
Improve collaboration
Networks companies
Marketing & customer
Community Networks Facebook/Reddit groups
engagement
Criminal Networks Terrorist or gang networks Law enforcement & intelligence

Dept of AI&DS, SIET Page 49


BUSINESS ANALYTICS BAD714B

Type Example Business Use


R&D collaboration between Encourage idea sharing and
Innovation Networks
firms innovation

SNA Metrics
SNA uses several quantitative metrics to analyze networks:
A. Connections-Based Metrics
Metric Meaning
Homophily People tend to connect with similar individuals
Strength increases when ties have multiple roles (e.g.,
Multiplexity
coworker + friend)
Reciprocity Mutual relationships (both follow each other)
Network Closure /
Your friends are also friends with each other
Transitivity
People connect with geographically or socially close
Propinquity
individuals
B. Distribution-Based Metrics
Metric Meaning Use
Centrality (Degree, Closeness, Who is most important in the
Find influencers
Betweenness) network
A person connecting two Useful for spreading
Bridge
groups messages
How interconnected the
Density Shows cohesion
network is
Strong vs. weak relationship Weak ties spread new
Tie Strength
bonds ideas faster
C. Segmentation-Based Metrics
Metric Meaning
Cliques Groups where everyone is connected to everyone
Clustering Coefficient Likelihood of group formation
Cohesion Minimum people needed to break the network

Social Media Analytics


Social media consists of platforms where users create and share content (e.g., Facebook,
Twitter, Instagram, YouTube, LinkedIn).
Types of Social Media (Kaplan & Haenlein, 2010)
Category Example Purpose
Collaborative Projects Wikipedia Shared knowledge
Blogs & Microblogs Twitter Personal opinion sharing
Content Communities YouTube Sharing media

Dept of AI&DS, SIET Page 50


BUSINESS ANALYTICS BAD714B

Category Example Purpose


Social Networking Sites Facebook, LinkedIn Connecting & communicating
Virtual Game Worlds World of Warcraft Immersive gaming
Virtual Social Worlds Second Life Digital identity interaction

FIGURE: Evolution of Social Media User Engagement.

Characteristics of Social Media vs. Traditional Media


Feature Traditional Media Social Media
Cost High Low
Reach Wide but controlled Wide and decentralized
Update Speed Slow Instant
Quality Control Edited and filtered Wide range (high to poor)
User Role Consumer only Producer + Consumer (Prosumer)
Interactivity Limited Highly interactive

Social Media Analytics Process


Companies analyze:
 Comments
 Likes/Shares
 Reviews
 Hashtags
 Follower interactions
To determine:
 Customer sentiment
 Public opinion
 Brand perception
 Competitor comparison
Types of Analytics Used
Type Explanation Example
Descriptive Analytics Tracks basic activity Number of followers, likes

Dept of AI&DS, SIET Page 51


BUSINESS ANALYTICS BAD714B

Type Explanation Example


Social Network Analytics Identifies influencers Who spreads trends?
Text & Sentiment Analysis Determines emotional tone “Love this product!” → Positive

Best Practices in Social Media Analytics


1. Use analytics for learning, not just scoring
→ Identify what works and improve continuously.
2. Track sentiment carefully
→ Separate positive and negative points in each message.
3. Continuously update text analysis rules
→ As slang, brand names, and language evolve.
4. Measure the "ripple effect"
→ Track how content spreads across networks.
5. Analyze beyond your brand
→ Understand the whole conversation about the market.
6. Find your true influencers
→ Influence ≠ number of followers, but power to change opinion.
7. Integrate insights into planning
→ Use analytics to guide marketing, customer service, and product decisions.

Social Analytics combines:


 Social Network Analysis → Who interacts with whom and how
 Social Media Analytics → What people are saying and how they feel
Together, they help organizations:
 Understand customer behavior
 Improve marketing strategies
 Strengthen brand reputation
 Predict trends and opportunities

Application Case 5.8: Tito’s Vodka Establishes Brand Loyalty with an Authentic Social
Strategy
If Tito’s Handmade Vodka had to identify a single social media metric that most accurately
reflects its mission, it would be engagement. Connecting with vodka lovers in an inclusive,
authentic way is something Tito’s takes very seriously, and the brand’s social strategy reflects
that vision. Founded nearly two decades ago, the brand credits the advent of social media
with playing an integral role in engaging fans and raising brand awareness. In an interview
with Entrepreneur, founder Bert “Tito” Beveridge credited social media for enabling Tito’s to
compete for shelf space with more established liquor brands. “Social media is a great
platform for a word-of-mouth brand, because it’s not just about who has the biggest
megaphone,” Beveridge told Entrepreneur. As Tito’s has matured, the social team has
remained true to the brand’s founding values and actively uses Twitter and Instagram to have
one-onone conversations and connect with brand enthusiasts. “We never viewed social media
as another way to advertise,” said Katy Gelhausen, Web & Social Media Coordinator. “We’re
on social so our customers can talk to us.” To that end, Tito’s uses Sprout Social to

Dept of AI&DS, SIET Page 52


BUSINESS ANALYTICS BAD714B

understand the industry atmosphere, develop a consistent social brand, and create a dialogue
with its audience. Recently and as a result, Tito’s organically grew its Twitter and Instagram
communities by 43.5% and 12.6%, respectively, within 4 months.
Informing a Seasonal, Integrated Marketing Strategy
Tito’s quarterly cocktail program is a key part of the brand’s integrated marketing strategy.
Each quarter, a cocktail recipe is developed and distributed through Tito’s online and offline
marketing initiatives. It is important for Tito’s to ensure the recipe is aligned with the brand’s
focus as well as larger industry direction. Therefore, Gelhausen uses Sprout’s Brand
Keywords to monitor industry trends and cocktail flavor profiles. “Sprout has been a really
important tool for social monitoring. The Inbox is a nice way to keep on top of hashtags and
see general trends in one stream,” said Gelhausen. These learnings are presented to Tito’s in-
house mixology team and used to ensure the same quarterly recipe is communicated to the
brand’s sales team and across marketing channels. “Whether you’re drinking Tito’s at a bar,
buying it from a liquor store or following us on social media you’re getting the same
quarterly cocktail,” said Gelhausen. The program ensures that, at every consumer touchpoint,
a person is receiving a consistent brand experience—and that consistency is vital. In fact,
according to an Infosys study on the omnichannel shopping experience, 34% of consumers
attribute cross-channel consistency as a reason they spend more with a brand. Meanwhile,
39% cite inconsistency as a reason enough to spend less. At Tito’s, gathering industry
insights starts with social monitoring on Twitter and Instagram through Sprout. But the
brand’s social strategy doesn’t stop there. Staying true to its roots, Tito’s uses the platform on
a daily basis to authentically connect with customers. Sprout’s Smart Inbox displays Tito’s
Twitter and Instagram accounts in a single, cohesive feed. This helps Gelhausen manage
inbound messages and quickly identify which require a response. “Sprout allows us to stay on
top of the conversations we’re having with our followers. I love how you can easily interact
with content from multiple accounts in one place,” said Gelhausen.
Spreading the Word on Twitter
Tito’s approach to Twitter is simple: engage in personal, one-on-one conversations with fans.
Dialogue is a driving force for the brand, and over the course of 4 months, 88% of Tweets
sent were replies to inbound messages. Using Twitter as an open line of communication
between Tito’s and its fans resulted in a 162.2% increase in engagement and a 43.5% gain in
followers. Even more impressively, Tito’s ended the quarter with 538,306 organic
impressions—an 81% rise. A similar strategy is applied to Instagram, which Tito’s uses to
strengthen and foster a relationship with fans by publishing photos and videos of new recipe
ideas, brand events and initiatives.
Capturing the Party on Instagram
On Instagram, Tito’s primarily publishes lifestyle content and encourages followers to
incorporate the brand into everyday occasions. Tito’s also uses the platform to promote its
cause marketing efforts and to tell its brand story. The team finds value in Sprout’s Instagram
Profiles Report, which helps them identify what media is receiving the most engagement,
analyze audience demographics and growth, dive deeper into publishing patterns, and
quantify outbound hashtag performance. “Given Instagram’s new personalized feed, it’s
important that we pay attention to what really does resonate,” said Gelhausen. Using the
Instagram Profiles Report, Tito’s has been able to measure the impact of its Instagram

Dept of AI&DS, SIET Page 53


BUSINESS ANALYTICS BAD714B

marketing strategy and revise its approach accordingly. By utilizing the network as another
way to engage with fans, the brand has steadily grown its organic audience. In 4 months,
@TitosVodka saw a 12.6% rise in followers and a 37.1% increase in engagement. On
average, each piece of published content gained 534 interactions, and mentions of the brand’s
hashtag, #titoshandmadevodka, grew by 33%.
Where to from Here?
Social is an ongoing investment in time and attention. Tito’s will continue the momentum the
brand experienced by segmenting each quarter into its own campaign. “We’re always getting
smarter with our social strategies and making sure that what we’re posting is relevant and
resonates,” said Gelhausen. Using social to connect with fans in a consistent, genuine, and
memorable way will remain a cornerstone of the brand’s digital marketing efforts. Using
Sprout’s suite of social media management tools, Tito’s will continue to foster a community
of loyalists. Highlights:
• A 162% increase in organic engagement on Twitter
• An 81% increase in organic Twitter impressions
• A 37% increase in engagement on Instagram

Questions for Discussion


1. How can social media analytics be used in the consumer products industry?
2. What do you think are the key challenges, potential solutions, and probable results in
applying social media analytics in consumer products and services firms?

Dept of AI&DS, SIET Page 54

You might also like