0% found this document useful (0 votes)
12 views50 pages

Text Analytics and Mining Explained

Text analytics and text mining are essential for organizations to convert unstructured text data into actionable insights, with text mining focusing on discovering patterns within that data. Natural Language Processing (NLP) enhances text mining by enabling computers to understand and process human language, addressing challenges such as ambiguity and context. Applications of text mining span various fields including marketing, security, and biomedicine, providing significant advantages in decision-making and knowledge extraction.

Uploaded by

hemanthv934
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views50 pages

Text Analytics and Mining Explained

Text analytics and text mining are essential for organizations to convert unstructured text data into actionable insights, with text mining focusing on discovering patterns within that data. Natural Language Processing (NLP) enhances text mining by enabling computers to understand and process human language, addressing challenges such as ambiguity and context. Applications of text mining span various fields including marketing, security, and biomedicine, providing significant advantages in decision-making and knowledge extraction.

Uploaded by

hemanthv934
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Text Analytics and Text Mining Overview

In today’s information age, organizations collect massive amounts of data, and most of it—
about 85%—is unstructured text (emails, reports, social media, documents). These
unstructured text data are growing rapidly, doubling every 18 months. Since business
decisions depend on converting data into knowledge, companies that can analyze their text
data effectively gain a strategic advantage.
Text analytics and text mining help transform unstructured text into meaningful and
actionable information using techniques such as natural language processing (NLP).

● Text Analytics is a broad field. It includes:

 Searching and retrieving relevant documents


 Extracting key information from text
 Applying data mining and web mining techniques

● Text Mining is more specific. It focuses on:

 Discovering new patterns and knowledge hidden in text data


Thus, text mining is considered a part of the larger area of text analytics.
Difference Between Text Analytics and Text Mining

● Text Analytics = Information Retrieval + Information Extraction + Data Mining +


Web Mining
(or simply Text Analytics = Information Retrieval + Text Mining)

● Usage Context:

o Text Analytics term commonly used in business

o Text Mining
term commonly used in academic and research environments

Even though definitions differ slightly, both terms are often used interchangeably.

What is Text Mining?


Text mining (also called text data mining or knowledge discovery in text) is a semi-
automated process of finding useful patterns and knowledge from large collections of
unstructured text.
It follows the same goal as data mining — discover valid, new, useful, and understandable
patterns, but:

Data Mining Text Mining

Works on structured data (tables, Works on unstructured text (PDFs, Word docs,
databases) XML, emails)

Two major steps in text mining:


1. Convert unstructured text into structured data.
2. Apply data mining techniques to extract patterns and knowledge.
Where is Text Mining Applied?
Useful in fields that generate huge text data, such as:

● Law (court records)

● Finance (annual/quarterly reports)

● Research (papers/articles)

● Medicine (hospital discharge summaries)


● Marketing (customer feedback, social media)

● Technology (patent files)

Example:
Using customer complaints/reviews to identify product defects or improve services.

Common Applications of Text Mining

Application What it does

Information
Find key phrases or relationships in text.
Extraction

Topic Tracking Recommends documents based on user interest.

Summarization Creates a shorter version of a long document.

Categorization Assigns documents to predefined categories.

Groups documents based on similarity (no predefined


Clustering
labels).

Concept Linking Connects documents with common ideas or themes.

Question Answering Automatically finds answers based on text content.

6.3 Natural Language Processing (NLP)


Earlier text mining systems used a simple method called the bag-of-words, where documents
are treated as unordered collections of words, ignoring grammar and meaning. This works for
simple tasks like spam detection but does not capture the real meaning of text. In many
cases (e.g., medical research papers), bag-of-words fails to give accurate results, showing the
need for more advanced techniques.
What is NLP?
Natural Language Processing (NLP) is a field of Artificial Intelligence that enables
computers to understand and process human language.
NLP aims to go beyond word counting and consider:

● Grammar

● Meaning (semantics)

● Context
NLP converts text into structured, machine-understandable data.

Challenges in NLP
Understanding language is difficult because human language is complex, ambiguous, and
context-dependent. Key challenges include:

NLP Challenge Explanation

Part-of-speech tagging Same word can be noun/verb depending on context.

Some languages (like Chinese) don’t have spaces between


Text segmentation
words.

Word sense One word can have multiple meanings; context decides
disambiguation meaning.

Syntactic ambiguity A sentence can be interpreted in multiple ways.

Imperfect input Errors, accents, and informal text complicate understanding.

Words may imply actions (e.g., “Pass the salt” is not a yes/no
Speech acts
question).

NLP Benefits and Applications


NLP helps computers automatically extract knowledge from text.
Stanford researchers enhanced WordNet, a large database of English words and semantic
relationships, using NLP. This helps build more intelligent language models.
One impactful business application is:

● Sentiment Analysis (measuring customer opinions from reviews, social media posts,
etc.)
Companies use NLP-based analytics to understand customer emotions and improve products
and services. NLP has successfully been applied to a variety of domains for a wide range of
tasks via computer programs to automatically process natural human language that previously
could only be done by humans.
Following are among the most popular of these tasks:
• Question answering. The task of automatically answering a question posed in
natural language; that is, producing a human language answer when given a human
language question. To find the answer to a question, the computer program may
use either a prestructured database or a collection of natural language documents (a
text corpus such as the World Wide Web).
• Automatic summarization. The creation of a shortened version of a textual
document by a computer program that contains the most important points of the
original document.
• Natural language generation. Systems convert information from computer databases
into readable human language.
• Natural language understanding. Systems convert samples of human language
into more formal representations that are easier for computer programs to
manipulate.
• Machine translation. The automatic translation of one human language to another.
• Foreign language reading. A computer program that assists a nonnative language
speaker to read a foreign language with correct pronunciation and accents
on different parts of the words.
• Foreign language writing. A computer program that assists a nonnative language
user in writing in a foreign language.
• Speech recognition. Converts spoken words to machine-readable input. Given a
sound clip of a person speaking, the system produces a text dictation.
• Text-to-speech. Also called speech synthesis, a computer program automatically
converts normal language text into human speech.
• Text proofing. A computer program reads a proof copy of a text to detect and
correct any errors.
• Optical character recognition. The automatic translation of images of handwritten,
typewritten, or printed text (usually captured by a scanner) into machine-editable
textual documents.

6.4 Text Mining Applications


Why Text Mining?

● Organizations collect huge amounts of unstructured data (emails, reports, customer


chats, reviews).
● Text mining helps convert this unstructured text into meaningful knowledge for
decision-making.

1. Marketing Applications

● Companies analyze customer conversations (call center notes, chat transcripts) using
text mining.

● Helps understand customer sentiments, complaints, satisfaction, etc.

● Text from blogs, reviews, discussion forums gives insights into public opinion about
products.

● Helps improve:

 Cross-selling & up-selling (suggesting related products to customers)


 Customer Relationship Management (CRM)

● Used to predict customer churn (customers likely to leave) so companies can take
action to retain them.

● Can automatically extract product attributes from product descriptions on websites;


useful for:
 Product recommendations
 Demand forecasting
 Comparing products across retailers
2. Security Applications

● Used for surveillance and intelligence gathering.

● ECHELON system (highly classified) is believed to intercept and analyze global


communications (emails, calls, fax, etc.).

● EUROPOL's OASIS system integrates data/text mining to track transnational


organized crime.
● The FBI and CIA, earlier working in separate databases, are now developing a
combined massive data warehouse with text mining modules to support law
enforcement across different levels.
3. Deception Detection (Truth vs. Lie Identification)

● Text mining can detect whether a written statement is truthful or deceptive.

● Researchers analyzed criminal statements and built a model that predicted deception
with 70% accuracy.

● Key advantages:

● Uses only text (no voice tone or body language needed).

● More practical and less intrusive compared to techniques like the polygraph.

Biomedical Applications of Text Mining


Why Text Mining is Powerful in Biomedicine

● Biomedical research produces huge volumes of literature (research papers, journals,


reports).

● Medical terminology is standardized and consistent, making it easier to extract


information.

● Text mining helps researchers find useful knowledge quickly from thousands of
documents.
How Text Mining Helps in Biomedical Research
1. Analyzing large experimental datasets
 Techniques like DNA microarrays, SAGE, and proteomics generate massive
gene/protein data.
 Text mining helps scientists compare this new experimental data with older
published research.
 Saves time during experiment validation and interpretation.
2. Predicting protein locations inside cells
 Protein location reveals its biological role and potential as a drug target.
 Shatkay et al. (2007) developed a system that uses:

▪ Text-based features (from research papers)


▪ Sequence-based features (from protein data)

 Their system outperformed earlier models in predicting protein location.


3. Finding disease–gene relationships
 Chun et al. (2006) created a system that scans MEDLINE to extract
relationships between diseases and genes.
 Uses:

▪ A dictionary of disease/gene names from public databases

▪ Machine learning (Named Entity Recognition – NER) to remove false


matches
 Result: Improved accuracy by 26.7% in identifying true relationships.
4. Extracting gene–protein / protein–protein interactions from literature
 Text mining can analyze biomedical sentences to identify interactions.
 Process:

▪ The text is tokenized (split into words)

▪ Part-of-speech tagging and shallow parsing are applied

▪ Words are matched against a domain ontology (hierarchical biomedical


knowledge base)
 Helps derive relationships between genes and proteins.
5. Contribution to large scientific initiatives
 This text mining method (Nakov et al., 2005) helps decode biological
relationships.
 Offers great potential to understand complex data in projects like the Human
Genome Project.

Academic Applications of Text Mining


1. Publishers use text mining to improve information retrieval
 Scientific publishers store huge databases of research papers.
 Text mining helps index and organize this information so researchers can
find specific content quickly.
2. Initiatives to support text mining in academic publishing
 Nature proposed an Open Text Mining Interface so computers can extract
meaning from articles.
 The National Institutes of Health (NIH) created a standard structure called
Journal Publishing Document Type Definition.
 These initiatives give machines semantic cues (meaning-based hints) to
answer queries from texts, without violating publishing restrictions.
3. Universities and research centers are adopting text mining
 The National Centre for Text Mining (NaCTeM) in the UK (University of
Manchester + University of Liverpool):

▪ Offers customized tools and research support.

▪ Initially focused on biomedicine; now expanded to social sciences.

4. Academic projects supporting scientific research


 The University of California, Berkeley is developing a project called BioText.
 It helps bioscience researchers perform text mining and analyze large sets of
scientific literature.
6.5 Text Mining Process
1. Need for a Standard Process
 Like data mining uses CRISP-DM, text mining also needs a structured
methodology.
 Text mining projects require more complex preprocessing because data is
unstructured (text, sentences, documents).
2. Context Diagram Explanation (High-Level View)
 Shows what is included and excluded in the text mining process.
3. Components of the Text Mining Process
 Input (left side):

▪ Text documents (unstructured data)

▪ Sometimes structured data (tables, databases)

 Output (right side):

▪ Knowledge or insights that help in decision-making

 Controls / Constraints (top):

▪ Software and hardware limitations

▪ Privacy issues
▪ Natural language complexity (difficulties in understanding human
language)
 Mechanisms / Resources (bottom):

▪ Proper techniques and algorithms

▪ Text mining software tools

▪ Domain experts (people with knowledge of the field)

4. Goal of Text Mining


 Convert unstructured text into meaningful, actionable knowledge.

5. Three Major Tasks (high-level process)
 Though detailed steps are not yet shown here, the process generally includes:
0. Preprocessing (prepare and clean the text)
1. Text Mining / Pattern Extraction
2. Evaluation & Interpretation (turn patterns into useful knowledge)

Task 1: Establish the Corpus (Collect Documents)


● Gather all documents related to the topic or domain.

● These documents can be in different formats:

 Text files, emails, web pages, XML files, notes


 Voice recordings can be converted into text using speech recognition

● After collecting, convert all documents into a common format (like plain text) so the
computer can process them.

● Documents can be stored:

 In folders as text files


 As links to web pages

● Text mining software can take these files and prepare them for further analysis.

Task 2: Create the Term–Document Matrix (TDM)

● A Term–Document Matrix (TDM) is created from the collected documents.

 Rows = documents
 Columns = terms (words)
 Cell values = how many times a term appears in a document

● Not all words are useful, so unnecessary terms are removed:

 Stop Words (e.g., “the”, “is”, “a”) — common words with no meaning for
analysis
 Synonyms can be grouped together (e.g., “car” and “automobile”)
 Phrases are treated as a single term (e.g., “Eiffel Tower”)

● Stemming is used:

 Converts words to their root form (e.g., modeling, modeled model)

● After cleaning, the TDM becomes smaller and more meaningful.


Representing the Indices (Normalization)

● After creating the Term–Document Matrix (TDM), each cell shows how many times a
word appears in a document.


Higher frequency always more important.

● Therefore, raw word counts need to be normalized to make comparisons more


meaningful.

● Common normalization methods:

 Log frequency (reduces the effect of very high word counts)


 Binary frequency (just marks whether word appears or not: 0 or 1)
 Inverse Document Frequency (IDF) (gives more weight to rare but
important words)
Reducing the Dimensionality of the Matrix

● TDMs are usually very large and sparse (most cells are zero).

● Goal: Reduce the size of the matrix to make analysis easier and faster.

● Methods to reduce size:

1. Manual filtering – domain expert removes irrelevant terms.


2. Remove rare terms – eliminate words that appear very few times.
3. Apply SVD (Singular Value Decomposition) – mathematical method to
reduce dimensions.

● SVD:

 Similar to Principal Component Analysis (PCA)


 Reduces many terms into fewer "concept dimensions"
 Helps reveal hidden meanings (latent concepts/semantic relationships)

Task 3: Extract the Knowledge


Knowledge Extraction Methods
Main techniques used to find patterns from structured data (like TDM):
[Link]
[Link]
[Link]
4. Trend Analysis
1. Classification

● Goal: Assign data (or text) to predefined categories or classes.

● Examples: Spam filtering, web page categorization, automatic tagging.

● Approaches:

 Knowledge Engineering: Uses expert-defined rules.


 Machine Learning: Learns automatically from labeled examples (now more
popular).

● In Text Mining: Called text categorization, used to find the right topic for each
document.
2. Clustering

● Goal: Group similar (unlabeled) items into natural clusters.

● Type: Unsupervised learning (no predefined labels).

● Use Cases: Document organization, web content grouping, search optimization.

● Benefits:

 Improved recall: Finds related documents beyond single-term matches.


 Improved precision: Groups related results for easier browsing.

● Popular Methods:

 Scatter/Gather Clustering: Dynamically groups documents to aid browsing.


 Query-Specific Clustering: Builds hierarchies of relevant documents for
better search relevance.
3. Association

● Meaning: Finds relationships or patterns between concepts/terms in large text


datasets.

● Goal: Discover which items or ideas often occur together.

● Key Measures:

 Support: % of documents containing both concept sets (A and C).


 Confidence: % of documents with A that also contain C.

● Example:

 “Software Implementation Failure” often appears with “ERP” and “CRM”


 Support = 4%, Confidence = 55%.

● Use Case:

 Used to detect links between ideas in literature, e.g., tracking the spread of
bird flu through related terms (regions, species, treatments).
4. Trend Analysis

● Meaning: Studies how concept frequencies change over time or across document
collections.

● Goal: Identify evolving topics, interests, or emerging trends.

● Example:

 Comparing research papers from different years to see how key ideas in
information systems have developed.

● Use Case:

 Helps in understanding topic evolution in academic or news datasets.


6.6 Sentiment Analysis & Topic Modeling

● Sentiment Analysis identifies opinions or emotions expressed in text (e.g., positive,


negative, neutral).

● Topic Modeling automatically discovers the main themes or topics present in a large
collection of documents.

● Both techniques are widely used because organizations generate large volumes of text
(emails, reviews, social media posts, etc.).

● These methods help businesses and researchers understand what people are talking
about and how they feel about those topics.
Sentiment Analysis

● People rely on others’ opinions (online reviews, social media posts, blogs) before
making decisions (e.g., buying a car, choosing a restaurant).

● Sentiment analysis (also called opinion mining) automatically identifies opinions and
emotions in text.

● It tries to answer the question: “What do people feel about a certain topic?”

● It classifies text usually into positive or negative, sometimes with levels (e.g., star
ratings).

● Sentiment can be:


 Explicit clearly expressed opinion (e.g., “This product is amazing.”)

 Implicit opinion implied indirectly (e.g., “The handle broke easily.”)


Used in many fields especially business, marketing, and customer relationship management.

● Helps companies analyze customer feedback from websites, reviews, tweets, etc.

● Real-time sentiment analysis enables companies to understand customers quickly and


take action.
Sentiment Analysis Applications
Sentiment analysis uses text from online sources to understand people's opinions and
feelings. It is widely used in business, marketing, finance, politics, and more.
Key Application Areas:
1. Voice of the Customer (VOC)
 Understands customer feedback from reviews, complaints, and social media.
 Helps improve products/service and react quickly to negative comments.
2. Voice of the Market (VOM)
 Looks at overall market trends and competitor opinions.
 Useful for market research and product positioning.
3. Voice of the Employee (VOE)
 Analyzes employee opinions from surveys, emails, and internal messages.
 Helps improve work culture and employee satisfaction.
4. Brand Management
 Monitors social media for brand reputation.
 Identifies threats (negative buzz) and opportunities (positive feedback).
5. Financial Markets
 Tracks market sentiment (news, blogs, tweets) to predict stock movements.
 Helps traders make informed decisions.
6. Politics
 Analyzes public opinion before elections.
 Helps understand voter concerns and predict election results.
7. Government Intelligence
 Detects hostile or negative communication spikes.
 Helps in security and monitoring potential threats.
8. Other Uses
 Improves e-commerce product suggestions.
 Filters emails (e.g., urgent negative emails go to priority folder).
 Opinion-based search engines (summarizing reviews).

Sentiment Analysis Steps


Step 1: Sentiment Detection (Objective vs. Subjective)

● First, check if the text contains an opinion or just a fact.


Text with high objectivity no opinion skip.
● Usually done by checking adjectives (e.g., “wonderful,” “bad”).

● Output = Objectivity–Subjectivity score (0 to 1).

Step 2: N–P Polarity Classification (Negative–Positive)

● For subjective text, decide whether the sentiment is positive or negative.

● Strength can also be identified:

 Mildly positive

 Moderately positive

 Strongly negative, etc.

● Sometimes texts contain mixed emotions find the dominant sentiment.

● Polarity can be analyzed at multiple levels:



Word Phrase Sentence Document
(one level’s output can feed the next).

Step 3: Target Identification

● Identify what the sentiment is about:

○ A product, person, event, etc.

● Easy for product reviews (target is obvious).

● Hard for general texts (news, blogs) with many possible targets.

● Comparative sentences (“A is better than B”) require identifying multiple targets and
ordering them.

Step 4: Collection and Aggregation

● Combine all individual sentiment values into one overall sentiment for the entire
document.

● Can be simple (sum/average) or complex (semantic NLP methods).


Polarity Identification Methods
Polarity (positive/negative/objective) can be identified at different levels:

● Word

● Term/Phrase

● Sentence

● Document

The most detailed level is word-level, and higher levels are formed by aggregating word
sentiments.

There are two main techniques for identifying polarity:

1. Using a Lexicon (Dictionary-Based Method)

A lexicon = a list of words + their meanings + sometimes their sentiment scores.

How it works Using a Collection of Training Documents

● Each word has a predefined positive, negative, or objective score.

● These scores are used to determine sentiment for larger text units.

Examples
● WordNet: A large dictionary grouped into synonym sets (synsets).

● SentiWordNet: An extension of WordNet where each synset has:

 Positivity score

 Negativity score

 Objectivity score

● WordNet-Affect: Adds emotion or feeling labels (anger, joy, fear, etc.).

How these lexicons are built

● Some are built manually.


● Others use algorithms to label words based on:

 Synonyms

 Antonyms

 Seed words (e.g., “good”, “bad”, “like”, “hate”).

2. Using Training Documents (Corpus-Based / Machine-Learning Method)

How it works

● Use a collection of human-labeled documents (e.g., movie reviews with


positive/negative tags).

● Machine learning identifies which words tend to appear in:

 Positive texts

 Negative texts

● A model learns to predict polarity based on patterns.

Benefit

● Domain-specific (e.g., works better for product reviews, news, social media).

Using Training Documents

1. Using a Collection of Training Documents (Supervised Learning)


What it is

● Uses manually labeled text (e.g., star ratings on Amazon, IMDb, etc.).

● These labels show whether the review is positive or negative.

● This makes it a supervised learning method (because training data has labels).

Where training data comes from

● Websites: Amazon, eBay, CNET, Rotten Tomatoes, IMDb.


● Research datasets from:

○ TREC

○ NII Test Collections

○ CLEF

● Custom datasets by researchers.

Machine learning algorithms used

● Neural Networks

● Support Vector Machines (SVM)

● k-NN

● Naive Bayes

● Decision Trees

● EM-based Clustering

2. Identifying Semantic Orientation of Sentences & Phrases

Goal

● Decide if a sentence or phrase is positive or negative.

How it's done

● Start from word-level polarity.

● Combine/average word polarities to get sentence/phrase polarity.

● Sometimes use ML models for more complex relationships.

3. Identifying Semantic Orientation of Documents

Goal

● Give a single sentiment label to the whole document.


How it's done

● Usually by averaging sentence/phrase scores.

● Works best for short/medium documents (like reviews).

● Not meaningful for very long documents.

4. Topic Modeling (Topic Detection)

What it is

● A way to automatically find themes/topics inside a large collection of documents.

● Assumes:

1. Every document has multiple topics.

2. Every topic is made of certain words.

5. Early Topic Modeling Methods

a) Clustering

● Group documents based on word frequencies.

● Label clusters by the most frequent words.

● Simple but limited.

6. LSA/LSI (Latent Semantic Analysis/Indexing)

Steps

1. Create a document-term matrix (rows = documents; columns = words).

2. Use tf-idf values instead of raw counts for better accuracy.

3. Perform SVD (Singular Value Decomposition) to reduce dimensions:

○ A=U*S*V


This separates words topics documents.
4. Find similar documents using cosine similarity.

Limitations

● Older models assumed:

○ A document belongs to one single topic, which is not realistic.

7. Latent Dirichlet Allocation (LDA)

● Solves LSA’s limitation.

● Allows:

○ One document = multiple topics (with probabilities).

○ One word = can belong to multiple topics.

● More realistic and widely used.

Latent Dirichlet Allocation (LDA)

1. What is LDA?

● LDA is one of the most popular topic modeling techniques.

● It helps discover hidden topics inside a collection of documents.

● It is still widely used, even though newer techniques (like deep learning and
Word2Vec) exist.

2. Type of Learning

● LDA is unsupervised
no labeled data is needed.
● It is a generative probabilistic model
it assumes documents are created from hidden topics.
3. Core Concept

LDA uses Dirichlet distributions to model:

● Document Topic relationships

● Topic Word relationships

Dirichlet distribution

● A probability distribution used for modeling multiple outcomes.

● It is a generalized form of the Beta distribution (for multiple categories instead of


two).

4. How LDA Thinks About Documents

LDA assumes:

1. Each document contains a mixture of several topics (e.g., 20% politics, 40% science,
40% health).

2. Each topic
is made of several words with different probabilities (e.g., Topic = “sports” words like game, team, w

5. Why LDA Is Useful

● It automatically finds patterns and categories inside large text collections.

● Helps summarize big document sets.

● Works well for real-world applications (news, research papers, reviews, etc.).

6. Famous Work & Examples

● The main paper: Blei et al., 2003 (seminal LDA work).


● David Blei (creator) demonstrated LDA by:

○ Applying a 100-topic LDA model to 17,000 Science journal articles.

○ Showed how meaningful scientific topics emerged automatically.

6.7 WEB MINING OVERVIEW


1. Impact of the Internet on Business

 The Internet has transformed business forever.

 Companies now face:

o More opportunities (global customers, new markets)

o More challenges (tough global competition)

 Being online is no longer optional — customers expect products, services, and


support online.

 Customers also share their experiences on social media, influencing others.

2. Explosion of Data on the Web

 Internet technologies make creating and sharing data extremely easy.

 Everything from delays, service issues, customer opinions → is now public.

 Smart companies use online tools to:

o Communicate better

o Understand customer needs

o Improve services

3. The Web as a Huge Information Source

 The Web is the largest source of data and text in the world.

 Contains:

o Text (HTML, XML)

o Hyperlinks (connections between pages)

o Usage data (clicks, visits, logs)

 This information can be mined to improve websites and understand users.

4. Challenges in Web Mining

Mining the Web is hard because:

a) The Web is too big


 Too large to store or analyze everything.

 Grows constantly.

b) The Web is too complex

 Web pages have no fixed structure.

 Different styles, formats, layouts.

c) The Web changes constantly

 Content gets updated frequently → news, blogs, prices, weather, stock updates.

d) The Web covers every domain

 People with different languages, backgrounds, purposes use it.

 Hard to understand what each user wants.

e) The Web contains everything

 Most of the Web is irrelevant to any single person.

 The challenge is finding the small useful portion.

5. Why Search Engines Alone Are Not Enough

 Keyword searches return:

 Too many results (many irrelevant)

 Miss highly relevant pages that don’t use exact keywords

 Search engines cannot fully understand:

 Meaning

 Context

 Relationships between pages

6. What Is Web Mining?

Definition:

Web mining = discovering useful patterns, insights, or relationships from Web data.

Types of Web data:


1. Content (text, images, audio, video)

2. Structure (hyperlinks between pages)

3. Usage (clicks, navigation, browsing patterns)

Purpose:

 Turn large Web data into actionable knowledge for better decision-making.

7. Web Mining vs. Web Analytics

Web Analytics

 Focuses on: What happened on a website?

 Uses predefined metrics (visits, clicks, bounce rate)

 More descriptive in nature.

Web Mining

 Focuses on: Discovering hidden patterns and predicting behavior

 Uses advanced data mining and machine learning.

 More predictive and prescriptive.

Relationship

 Web analytics is a part of Web mining.

8. Three Main Types of Web Mining (Taxonomy)

1. Web Content Mining

 Extracting useful information from Web content


(text, videos, images, documents)

2. Web Structure Mining

 Analyzing hyperlinks between Web pages


(e.g., finding authoritative pages)

3. Web Usage Mining

 Analyzing user behavior and click patterns


(e.g., what users do on websites)
6.8 SEARCH ENGINES

Why Search Engines Are Important

 The web is huge and complicated; finding information manually is hard.

 People use search engines to:

 Compare products and prices.

 Read reviews and complaints.

 Explore places, people, events, services, etc.

 Search engines are central to most online activities.

 Google’s success shows how important they are.

What a Search Engine Does

 It is a program that finds documents/pages based on the words (keywords) typed by


the user.

 Search engines are also called information retrieval systems.

 Can be used on the web or on desktops/documents.

Goals of a Search Engine

 Effectiveness (quality): Find the right pages.

 Efficiency (speed): Show results fast.

 Improving one often reduces the other, so good engines balance both.

Anatomy of a Search Engine

A search engine has two main cycles:

1. Development Cycle (Backend / Production side)

Purpose: Build and organize a large database of web pages so results can be returned
quickly.

It has two major components:

A. Web Crawler (Spider)

 Software that visits websites automatically.

 Collects/copies pages to store them for later use.


 Starts from a list of initial websites (called seeds).

 Finds all hyperlinks on each page and adds them to a list of pages to visit.

 Follows rules made by the search engine (e.g., which pages to visit first).

 Cannot download everything at once, so it must prioritize.

Why Crawlers Exist

 The web is too large to search in real time.

 Crawlers create a cached (stored) version of the web in the search engine’s database.

2. Responding Cycle (Frontend / User side)

 Takes a user’s query and searches the pre-built index.

 Returns the best-matching documents/pages.

 Shows a ranked list of results.

DOCUMENT INDEXER

STEP 1: Preprocessing

Purpose: Convert all pages into a standard, uniform format.

 Web pages come in many formats (text, images, links, etc.).

 The indexer separates different content types.

 Converts everything into a standard structure so it’s easier to process later.

STEP 2: Parsing the Documents

Purpose: Extract meaningful words/terms using text-mining and NLP techniques.

This step includes:


 Tokenization: Breaking sentences into words/terms.

 Spell-checking and fixing errors using dictionaries (lexicons).

 Removing stop words: Words that don’t help in search (like “and”, “the”, “is”).

 Stemming: Reducing words to their root form


(e.g., “playing”, “played”, “plays” → “play”).

 Handling synonyms/homonyms using resources like WordNet.

 Result: A clean list of important, meaningful terms for each document.

STEP 3: Creating the Term-by-Document Matrix

Purpose: Determine which words appear in which documents and how important they
are.

 A matrix is created where:

o Rows = words/terms

o Columns = documents

o Values = weight of the term in that document

 Weighting can be:

o Binary (1/0) — term is present or not.

o Term Frequency (TF) — how many times the word appears in the document.

o TF-IDF — best method; increases importance of rare but meaningful terms


and lowers importance of common words.

Response Cycle of a Search Engine


The response cycle has two main parts:

1. Query Analyzer

2. Document Matcher/Ranker

1. Query Analyzer

Purpose: Understand the user’s search query and convert it into a form that matches the
document index.

It performs tasks similar to the document indexer:

 Tokenization → breaks the query into words.

 Removing stop words (like “the”, “is”, “and”).


 Stemming → reduces words to their root form.

 Spelling correction & synonym handling.

 Converts the query into the same standardized structure used for indexed
documents.

Why?
So the user’s query and document index use the same format, making matching faster and
more accurate.

2. Document Matcher/Ranker

Purpose: Find the most relevant documents and rank them in the correct order.

What it does:

 Matches the processed query against the document database.

 Finds documents that best match the query.

 Ranks them in order of relevance/importance.

 Each search engine uses its own ranking algorithm (often secret).

Early methods:

 Simple keyword matching.

 Ranking based on number of matched terms + their weights.

 Results were not very accurate.

PageRank (Google, 1997)

 A new algorithm created by Google.

 Ranks web pages based on importance, not just keyword matching.

 Looks at:

 Which pages link to a page.

 How important those linking pages are.

Important:
PageRank is an addition to normal keyword-based matching, not a replacement.

Improving Search Results

 Search engines track user behaviour after showing results.


 If users click results lower in the list, it may mean:

o Ranking was not perfect.

 This feedback helps refine ranking rules over time.

Search Engine Optimization (SEO)

What is SEO?

 SEO means improving a website so it appears higher in unpaid (organic) search


engine results.

 Higher ranking = more visibility = more visitors.

 SEO focuses on:

 How search engines work

 What users search for

 Which keywords they type

 Which search engines users prefer

How SEO Works

Improving a website may involve:

 Editing content, HTML, and code to match important keywords.

 Removing anything that stops search engines from indexing the site.

 Increasing backlinks (links from other sites), which boosts credibility.

How Search Engines Used to Work

 Earlier, website owners submitted URLs manually to search engines.

 A crawler visited that page, extracted links, and indexed it.

Today:

 Search engines continuously crawl the web on their own.

 They automatically find, fetch, and index information.

Why Ranking Matters

 Being just indexed is not enough.


 A website must:

 Rank higher than competitors.

 Appear frequently in search results.

Methods to improve ranking:

 Cross-linking between pages within the same website.

 Writing content with popular keywords.

 Updating content regularly (makes crawlers come back).

 Using good metadata (title tags, descriptions).

 Using canonical URLs so links are not split across variations.

 Redirecting duplicate URLs to a single page.

Types of SEO Methods

SEO is divided into two categories:

1. White-Hat SEO (Approved, Ethical)

 Follows search engine guidelines.

 Creates content for users, not machines.

 Makes content accessible to crawlers without tricking them.

 Provides the same content to users and search engines.

 Lasts long and is safe.

2. Black-Hat SEO (Unapproved, Risky)

 Tries to manipulate rankings through tricks.

 Examples:

 Hiding text (invisible or off-screen)

 Cloaking (showing one page to users and another to crawlers)

 Search engines may:

 Penalize the site by lowering rankings.

 Remove the site completely.


 Example: In 2006, BMW Germany and Ricoh Germany were removed for bad
practices.

Challenges of SEO

 SEO results can be profitable, but:

 Search engines make frequent algorithm changes.

 Rankings can drop unexpectedly.

 Heavy dependence on Google is risky.

 Google made 500+ algorithm changes in 2010 (≈1.5 per day).

Because of this uncertainty, businesses:

1. Hire SEO companies.

2. Pay for sponsored (paid) search results.

3. Avoid over-dependence on search engine traffic.

Bottom Line for E-commerce

 Traffic alone is not success.

 The real goal: convert visitors into customers (sales).

6.9 Web Usage Mining (Web Analytics)

What is Web Usage Mining?

 Also called Web analytics.

 It means extracting useful information from data produced by website visits and
transactions.

 Helps understand how users behave on a website.

What Data Is Analyzed?

 Data collected by Web servers, such as:

 Pages users click

 Order of clicks (clickstream)

 Time spent on pages

 Search queries
 Downloads

 Purchases

Clickstream Analysis

 The study of the path users follows when clicking through a website.

 Reveals patterns in user behavior.

Examples:

 If 60% of people who search for “hotels in Maui” also earlier searched for “airfares
to Maui,”
→ The company can place targeted ads smartly.

 If 70% of software downloads happen between 7–11 PM,


→ Company can increase support staff and bandwidth at that time.

How the Knowledge Is Used

Insights from clickstream analysis help a company to:

1. Improve processes
(Example: better scheduling, better customer support)

2. Improve the website


(Example: better navigation, faster loading pages)

3. Increase customer value


(Example: showing more relevant ads/products to users)

Web Analytics Technologies

What is Web Analytics?

 Tools and methods used to measure, collect, analyze, and report Internet data.

 Helps understand and optimize Web usage.

 Not only used for tracking traffic but also for:

 E-business insights

 Market research

 Improving e-commerce effectiveness

 Measuring impact of advertising campaigns


What Web Analytics Helps With

 Measuring number of visitors and page views.

 Studying traffic patterns and popularity trends.

 Understanding how marketing campaigns affect website traffic.

Two Main Types of Web Analytics

1. Off-Site Web Analytics

 Measurement that happens outside your website.

 Tracks:

 Potential audience (how many people you can reach)

 Share of voice (how visible your brand is online)

 Internet buzz (opinions, reviews, social media comments)

2. On-Site Web Analytics (more common)

 Measures visitor behavior on your own website.

 Tracks:

 How users arrive (drivers)


 What actions they take (conversions)
 Which pages lead to purchases

 Used to:

 Measure website performance

 Improve marketing campaigns

 Compare results with KPIs

 Tools: Google Analytics (most popular), Yahoo!, Microsoft tools, and many new
ones.
Two Ways to Collect On-Site Data

1. Server Log File Analysis

 Web server automatically records all page/file requests.

 Traditional method.

2. Page Tagging

 Uses JavaScript code on each webpage.

 Sends information to an analytics server when a page loads or a click happens.

 Popular method used by Google Analytics.

Additional Data Sources

To better understand user behavior, companies also combine:

 Email data

 Direct mail campaign results

 Sales and lead history

 Social media data

Web Analytics Metrics

Web analytics provides important data that helps businesses understand visitor behavior and
improve marketing decisions. The key actionable metric categories are:

1. Website Usability – How visitors use the site

These metrics show how user-friendly your website is and whether the content is effective.

Key Metrics:

 Page Views

 Shows average number of pages viewed per visitor.


 Low page views may mean poor design, confusing layout, or mismatched
marketing messages.

 Time on Site

 Shows how long visitors spend on your website.

 More time usually means users are interested and engaging.

 Compare with page views—long time + few pages may mean visitors can’t
find what they need.

 Downloads

 Measures how often resources like PDFs, videos, brochures are downloaded.

 Low downloads may mean poor visibility or poor promotion of these


resources.

 Click Map

 Shows where users click on the page (images, links, buttons).

 Helps check if important items are getting enough attention.

 Click Paths

 Tracks the path visitors follow through the website.

 Helps identify where users drop off or get stuck.

 Shows if users follow desired paths (e.g., learning → engaging → buying).

 Different processes for different users (beginners, returning users, buyers) can
be monitored.

Traffic Sources

Traffic sources show where your website visitors come from.

1. Referral Websites

 Visitors come from other websites that link to yours.

 Helps identify which referral sites bring:

 Most traffic

 Highest conversions

 Most new visitors

2. Search Engines (Paid & Organic)


 Shows which keywords people used to find your site.

 Helps check if the keywords match your products/services.

 Businesses may need many keywords since searches can vary a lot.

3. Direct Traffic

 Visitors come directly without another website.

 Happens when:

 They type your URL directly.

 They click a bookmark/favorite.

 They see your URL on offline materials (ads, brochures, radio).

 Coded URLs help track specific sources.

4. Offline Campaigns

 Print ads, TV, radio, brochures.

 Use a special URL (example: [Link]/offer50) to track how many people


respond to offline ads.

5. Online Campaigns

 Includes banner ads, Google ads, email campaigns.

 Each campaign can have a unique URL to measure performance.

Visitor Profiles

Visitor profiles help you understand who your visitors are and what they want.

1. Keywords

 The search terms people use reveal:

 Whether they already know your product

 Whether they are just looking for a solution

 Helps you tailor content for different types of visitors (new vs. informed).

2. Content Groupings

 Group website sections by product or campaign.


 Shows which sections get the most traffic based on promotions or trade shows.

3. Geography

 Shows where visitors come from: country, state, city.

 Useful for:

 Region-based marketing

 Geo-targeted ads

 Understanding market reach

4. Time of Day

 Shows when people visit your site (morning, lunch, evening).

 Helps decide:

 When people browse vs. when they buy

 Best hours for customer support

5. Landing Page Profiles

 Each advertising campaign can send visitors to its own landing page.

 Helps measure:

 Which audience group visited

 How many visitors came from each demographic or campaign

Conversion Statistics

A conversion means a visitor completed an action you want—like buying, registering, or


filling a form. Every company defines its own conversion goals.

1. New Visitors

 Shows how many people are coming to your website for the first time.

 Useful when your goal is increasing visibility and reach.

2. Returning Visitors

 Shows how many visitors come back.

 Important for:

 Loyalty programs
 Products with long decision-making cycles

3. Leads

 When a visitor submits a form → it becomes a lead.

 You can calculate:

 Completion rate = (completed forms ÷ total visitors who saw the form)

 Low completion means your form/page needs improvement.

4. Sales/Conversions

 A sale can be:

 An online purchase

 A registration

 A sign-up

 Any action you define as success

 Tracking this tells you what is working well in your website journey.

5. Abandonment / Exit Rates

 Abandonment: When users start a process (like checkout) but quit before finishing.

 Shows where users drop off.

 If many quit at the same step → something needs fixing there.

 Exit rate: When users leave a page quickly.

 High exit rate means:

 The page didn’t meet expectations.

 Your ads/messages may not match what the page delivers.

Why these metrics matter

 Together, these metrics help you:

 Identify strengths

 Spot problems

 Improve user experience

 Increase conversions and sales


 Many companies create a weekly dashboard to monitor all these numbers.

6.10 Social Analytics

1. Meaning of Social Analytics

 The term can mean different things, but in business/technology, it refers to:
Analyzing digital interactions and relationships between people, content, topics,
and ideas.

 It includes:

 Mining text from social media (e.g., sentiment analysis)

 Analyzing social networks (e.g., influence, connections)

2. Two Main Branches of Social Analytics

1. Social Network Analysis (SNA)

2. Social Media Analytics

Social Network Analysis (SNA)

What is a Social Network?

 A structure showing how people or groups are connected.

 These connections help identify:

 Who influences whom

 How information spreads

 Important members in the network

Purpose of SNA

 Understand patterns and relationships.

 Find influential people.

 Study how networks grow, function, and interact.

Origins

 Comes from sociology, psychology, statistics, and graph theory.

 Methods developed from the 1950s and 1980s.


Why SNA is Important Today

 It's used heavily in:

 Business analytics

 Consumer behavior

 Marketing

 Sociology

 Fraud detection

 Criminal investigations

Types of Social Networks Important for Business

1. Communication Networks

 Show how information flows between people.

 Useful for:

 Improving customer relations

 Understanding how customers share messages

 Telecom companies optimizing network usage

2. Community Networks

 Earlier: Based on geography (who interacts with whom locally).

 Now: Online communities formed on:

 Social media

 Forums

 Messaging apps

 Companies analyze these to:

 Discover customer interests

 Improve marketing

 Understand group behaviors

3. Criminal Networks
 Show relationships between criminals or gangs.

 Used by law enforcement to:

 Track illegal activities

 Understand gang structures

 Predict or prevent crimes

4. Innovation Networks

 Focus on how ideas and innovations spread.

 Used to understand:

 Why some groups adopt new ideas quickly

 How innovations move through communities

 Which networks are more creative or influential

Social Network Analysis Metrics

SNA studies networks made of nodes (people/organizations) and ties (connections).


The metrics are grouped into three categories:

1. Connections, 2) Distributions, and 3) Segmentation.

1. Connections Metrics

These describe how nodes are connected.

Homophily

 People connect more with others who are similar (same age, gender, interests, etc.).

Multiplexity

 When two people share multiple relationships.


Example: Two people are friends + coworkers → multiplexity = 2.

Mutuality / Reciprocity

 Measures if friendships/relationships are two-way.

Network Closure

 Measures if friends of a person are also friends with each other.


(Also called transitivity.)

Propinquity
 People tend to form ties with those who are physically/geographically close.

2. Distributions Metrics

These describe the structure and flow of the network.

Bridge

 A person who connects two groups that otherwise have no link.

 Weak ties often act as bridges → very important for spreading info.

Centrality

Shows how important a node is in a network.


Different types:

 Degree centrality – number of direct connections

 Betweenness centrality – how often a node lies on the shortest path between others

 Closeness centrality – how near a node is to all others

 Eigenvector centrality – importance based on being connected to other important


nodes

Density

 How connected the network is compared to how connected it could be.

Distance

 Minimum number of steps (ties) needed to reach from one person to another.

Structural Holes

 Gaps in the network where groups are not connected.

 Someone who fills the hole gains advantage (access to unique info).

Tie Strength

 Indicates relationship strength based on:

 time spent

 emotional closeness

 intimacy

 reciprocity
 Strong ties → close relationships

 Weak ties → good for new information, act as bridges

3. Segmentation Metrics

These identify groups or clusters within the network.

Cliques / Social Circles

 Clique: everyone is directly connected to everyone.

 Social circle: similar group but not all are directly connected.

Clustering Coefficient

 Measures how likely it is that two friends of a node are also friends.

 High value = network is more grouped/cliquish.

Cohesion

 Measures how tightly connected a group is.

 Structural cohesion: smallest number of people whose removal would break the
group apart.

Social Media

What is Social Media?

Platforms where people create, share, and exchange information, ideas, and opinions
online.

Based on Web 2.0, which allows user-generated content (posts, videos, comments,
etc.).

Uses mobile and web technologies to let people communicate, collaborate, and
interact.

Types of Social Media (Kaplan & Haenlein, 2010)

Collaborative projects – Wikipedia

Blogs & microblogs – Twitter

Content communities – YouTube

Social networking sites – Facebook


Virtual game worlds – World of Warcraft

Virtual social worlds – Second Life

Why Social Media is Different from Traditional Media

1. Quality

Traditional media = controlled, consistent quality (edited & approved).

Social media = quality varies a lot; can be excellent or very poor.

2. Reach

Both reach global audiences.

Traditional = centralized control (newspapers, TV).

Social media = decentralized, anyone can publish.

3. Frequency

Social media content can be updated quickly and repeatedly.

Traditional publishing is slower and costly.

4. Accessibility

Traditional media needs money and organizations to publish.

Social media is free or cheap, available to everyone.

5. Usability

Traditional media = needs professional skills.

Social media = anyone with basic skills can post.

6. Immediacy

Traditional media = long time to publish.

Social media = instant posting and feedback.

7. Updatability

Traditional = once printed, cannot be changed.

Social media = can modify or update anytime (edit, comment, repost).


Measuring Social Media Impact

Organizations analyze social media to understand customer opinions, trends, and the
impact of their marketing efforts.

Three Types of Social Media Analytics

1. Descriptive Analytics

 Measures basic numbers like:

 Followers count

 Number of reviews

 Most used channels

 Shows trends and activity levels.

2. Social Network Analysis

 Studies connections between people (followers, friends, influencers).

 Helps find:

 Who has the strongest influence

 How information spreads

3. Advanced Analytics

 Uses predictive and text analytics.

 Understands:

 Themes in conversations

 Sentiment (positive/negative/neutral)

 Hidden patterns

Usually, companies use all three for complete insights.

Social Media Analytics

1. Think of Measurement as Guidance, Not Judgment

 Analytics should guide improvement, not punish.

 Helps identify:

 What works
 What doesn’t

 Which platforms matter (Twitter vs Facebook)

2. Track Customer Sentiment Properly

 Understand emotional tone: positive, negative, or neutral.

 Don’t mix sentiments in one review.

 Example: “Great location but smelly bathroom”

 Separate into positive and negative.

 Compare sentiment with competitors.

3. Continually Improve Text Analysis

 Text analytics becomes better over time.

 Update rules, keywords, categories regularly.

 Tools learn business vocabulary gradually.

4. Study the Ripple Effect

 A post may:

 Go viral (retweeted/shared widely), OR

 Fade quickly

 Analytics should show what made it spread.

5. Look Beyond Your Brand

 Customers talk about:

 Experiences

 Needs

 Problems

 Not just your brand name.

 Track wider conversations related to your industry.

6. Identify Key Influencers

 Important influencers are:

 Not always brand promoters


 But those who can shape opinions in your field

 Understand what they say and how they impact your brand.

7. Check Accuracy of Analytics Tools

 Automated tools are improving but not perfect.

 Accuracy:

 Twitter/product reviews: 80–90%

 Blogs/forums: 60–70%

 Accuracy improves with better algorithms.

8. Use Social Media Insights in Planning

 Include social media findings in business decisions.

 Track activity spikes and connect them to:

 Campaigns

 Events

 Market changes

You might also like