0% found this document useful (0 votes)
10 views19 pages

UNIT2

The document provides an overview of text analytics and text mining, explaining their definitions, processes, techniques, and applications. Text analytics transforms unstructured text into structured insights, while text mining discovers hidden patterns and knowledge from large volumes of text. It also discusses the significance of IBM's Watson in demonstrating AI's capabilities in understanding and processing human language.

Uploaded by

akshithinavolu
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views19 pages

UNIT2

The document provides an overview of text analytics and text mining, explaining their definitions, processes, techniques, and applications. Text analytics transforms unstructured text into structured insights, while text mining discovers hidden patterns and knowledge from large volumes of text. It also discusses the significance of IBM's Watson in demonstrating AI's capabilities in understanding and processing human language.

Uploaded by

akshithinavolu
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

UNIT – II:

TEXT ANALYTICS AND TEXT MINING

1. Introduction to Text Analytics


Text analytics refers to the process of transforming unstructured text data into
meaningful and structured insights. It extracts patterns, keywords, sentiment,
and trends from sources like social media posts, reviews, emails, blogs, and
documents. Text analytics helps organizations understand customer opinions,
detect issues, and support decision-making.

2. Introduction to Text Mining


Text mining (or text data mining) is the technique of discovering hidden
patterns, relationships, and knowledge from large volumes of text. It
combines natural language processing (NLP), machine learning, and data
mining to extract useful information. It includes tasks like classification,
clustering, topic modeling, and sentiment analysis.

3. Differences Between Text Analytics and Text Mining

Text Analytics Text Mining

Focuses on analyzing and Focuses on extracting patterns and


interpreting text knowledge

More descriptive More exploratory and predictive

Discovers new information and


Produces insights and reports
relationships

4. Sources of Text Data


 Social media (Twitter, Facebook, Instagram)
 Emails and customer feedback
 Online reviews (Amazon, Google Reviews)
 Blogs and news articles
 Call center transcripts
 Survey responses

5. Text Analytics Process


1. Data Collection
2. Text Preprocessing
3. Tokenization
4. Stemming / Lemmatization
5. Feature Extraction (BoW, TF-IDF, Embeddings)
6. Modeling (Classification, Clustering, Sentiment Analysis)
7. Evaluation and Interpretation
8. Visualization (word clouds, charts)

6. Text Preprocessing Techniques

✔ a. Tokenization
Breaking text into smaller units (words/phrases).
✔ b. Stop-word Removal
Removing common words like the, is, of.
✔ c. Stemming
Reducing words to root form by removing suffixes.
Example: playing → play
✔ d. Lemmatization
Reducing words to dictionary form using grammar rules.
Example: better → good
✔ e. Normalization
Lowercasing, removing numbers, special characters.
7. Feature Extraction Methods

✔ Bag of Words (BoW)


Represents text by counting word frequency.
✔ TF-IDF
Assigns importance to words based on how unique they are.
✔ Word Embeddings
Advanced representation using vectors:
 Word2Vec
 GloVe
 BERT embeddings

8. Text Mining Techniques

✔ 1. Classification
Assigning text into categories (spam/not spam).
✔ 2. Clustering
Grouping similar documents together.
✔ 3. Sentiment Analysis
Understanding emotions: positive, negative, neutral.
✔ 4. Topic Modeling
Finding hidden topics in documents (LDA).
✔ 5. Information Extraction
Extracting names, dates, locations (NER).
✔ 6. Text Summarization
Generating short summaries of long documents.

9. Sentiment Analysis
Sentiment analysis identifies the emotional tone behind text.
Uses NLP to classify text into:
 Positive
 Negative
 Neutral
Applications: brand monitoring, customer feedback, social media analysis.

10. Applications of Text Analytics and Text Mining


 Social media monitoring
 Customer review analysis
 Fraud detection
 Healthcare diagnosis
 Legal document analysis
 Email filtering and spam detection
 Market research and trend detection
 Chatbots and customer support systems

11. Challenges in Text Analytics


 Language complexity
 Sarcasm and ambiguity
 Large unstructured text
 Domain-specific vocabulary
 Real-time processing
 Multilingual data
Machine Versus Men on Jeopardy – The Story of IBM Watson
“Machine Versus Men on Jeopardy” refers to the historic competition in 2011
where IBM’s AI system Watson competed against two of the greatest human
Jeopardy champions—Ken Jennings and Brad Rutter. The purpose was to
show how advanced analytics, machine learning, and natural language
processing (NLP) could enable computers to understand and answer complex
human language questions.

1. Background
Jeopardy is a quiz show known for its complex clues, which often include
jokes, puns, riddles, and indirect wording. Understanding these questions
requires not only knowledge but also language understanding—a major
challenge for machines.
IBM developed Watson, a supercomputer system designed to:
 Process natural language
 Search massive amounts of text
 Evaluate possible answers
 Choose the answer with the highest confidence

2. How Watson Worked


Watson used advanced techniques such as:
 Natural Language Processing (NLP) to understand human questions
 Text mining to search millions of documents
 Machine learning to rank possible answers
 Statistical models to choose the most confident response
It did not use the internet during the game. Instead, it relied on stored
encyclopedias, books, and datasets.

3. The Competition
Watson competed against:
 Ken Jennings – Longest winning streak
 Brad Rutter – Highest money winner
Watson answered questions faster and more accurately, eventually defeating
both human champions by a large score margin. This proved that machines
can process and interpret unstructured text at very high speeds.

4. Significance of Watson’s Victory


Watson’s success showed the power of:
 Big data analytics
 NLP and text mining
 Machine learning
 High-performance computing
It demonstrated how AI could go beyond simple calculations and understand
human language almost like people.

5. Real-World Applications After Jeopardy


Watson’s technology was later used in:
 Healthcare (diagnosis suggestions)
 Customer service (chatbots)
 Finance (risk analysis)
 Education (personalized learning)
 Business intelligence (decision support)

The Story of Watson (Exam Answer)


Watson is an advanced artificial intelligence (AI) system developed by IBM to
demonstrate how machines can understand and process human language. The
system became famous in 2011 when it competed on the American quiz show
Jeopardy! Against two of the greatest human champions, Ken Jennings and
Brad Rutter. Watson’s victory marked a major milestone in the evolution of AI,
natural language processing (NLP), and text analytics.
1. Why Watson Was Created
Watson was designed to show that a computer could understand natural
language, interpret complex questions, search large amounts of text, and deliver
accurate answers faster than humans. Jeopardy was chosen because its questions
involve:
 Riddles
 Wordplay
 Jokes
 Multi-layered clues
These are extremely difficult for machines to understand.

2. How Watson Worked


Watson used multiple advanced technologies:
 Natural Language Processing (NLP) to understand human questions
 Text Mining to analyze millions of documents
 Machine Learning to improve accuracy
 Deep Question Answering Algorithms to choose the best possible
answer
 Massive parallel computing to process information in seconds
Watson did not search the internet; it relied on a vast internal database of books,
encyclopedias, and articles.

3. Watson on Jeopardy!
In 2011, Watson competed against top human champions.
It:
 Interpreted the question
 Generated multiple possible answers
 Calculated a confidence score
 Buzzed in only when confident
In the end, Watson won the competition with a high score, proving that AI can
compete with humans in tasks involving language and knowledge.

4. Impact and Significance


Watson’s victory demonstrated that:
 Machines can understand and process unstructured text
 AI can support decision-making in real-world scenarios
 NLP and machine learning can outperform humans in specific tasks
Watson became a symbol of how AI can transform industries using data and
analytics.

5. Real-World Applications After Jeopardy


IBM expanded Watson’s technology into several fields:
 Healthcare: Helping doctors diagnose diseases and select treatments
 Finance: Risk analysis and investment insights
 Customer Service: Intelligent chatbots
 Business Analytics: Decision support and forecasting
 Education: Personalized learning support

Text Analytics and Text Mining – Concepts and Definitions

1. Text Analytics – Definition


Text Analytics is the process of transforming unstructured text data into
meaningful and structured insights. It involves analyzing text to extract patterns,
keywords, sentiments, topics, and useful information. Text analytics uses
techniques such as NLP, machine learning, and statistical analysis to support
decision-making.
2. Text Mining – Definition
Text Mining (also called Text Data Mining) is the process of discovering
hidden patterns, relationships, and knowledge from large volumes of text. It
goes deeper than text analytics by applying algorithms to identify trends,
classify documents, cluster similar text, and extract information automatically.

3. Concept of Text Analytics


The concept of text analytics is built on converting raw text into structured
data that computers can analyze. It includes cleaning text, breaking it into
components (tokens), removing noise, extracting features, and generating
insights. The goal is to understand human language data and convert it into
useful business intelligence.

4. Concept of Text Mining


The concept of text mining focuses on discovering new insights that are not
directly visible from the text. It involves exploring patterns, predicting
behaviors, finding relationships, grouping similar texts, detecting sentiment, and
identifying topics. It is widely used in social media, marketing, healthcare, fraud
detection, and more.

5. Key Techniques in Text Analytics & Text Mining

✔ a. Text Preprocessing
Includes cleaning, tokenization, stop-word removal, stemming, lemmatization.
✔ b. Feature Extraction
Converting text into numerical form using Bag of Words, TF-IDF, Word
Embeddings.
✔ c. Text Classification
Assigning text to categories (e.g., spam or not spam).
✔ d. Clustering
Grouping similar text documents (e.g., news articles).
✔ e. Sentiment Analysis
Identifying emotions in text (positive, negative, neutral).
✔ f. Topic Modeling
Finding hidden themes in text using methods like LDA.
✔ g. Named Entity Recognition (NER)
Extracting names of people, locations, dates, etc.

6. Applications of Text Analytics & Text Mining


 Social media monitoring
 Customer feedback analysis
 E-commerce product review analysis
 Fraud and spam detection
 Healthcare clinical text analysis
 Email filtering
 Market intelligence and trend detection
 Chatbots and customer service

7. Benefits
 Helps organizations understand customer opinions
 Supports decision-making with insights
 Converts unstructured text into actionable knowledge
 Saves time by automating text analysis
 Improves marketing, product design, and customer service
Natural Language Processing (NLP) – Definition
Natural Language Processing (NLP) is a field of Artificial Intelligence (AI)
that enables computers to understand, interpret, and generate human language.
NLP helps machines read text, hear speech, understand sentences, extract
meaning, and respond in a way that is meaningful to humans.
In simple words:
NLP teaches computers to understand human language like English, Hindi,
etc.

Concept of NLP
The concept of NLP is to bridge the gap between human communication and
computer understanding. It combines:
 Linguistics (grammar, meaning, sentence structure)
 Computer science
 Machine learning
NLP converts unstructured text or speech into structured information that
machines can analyze and use for decision-making.

Key Tasks in NLP


1. Tokenization – Splitting text into words or sentences.
2. Part-of-Speech Tagging – Identifying nouns, verbs, adjectives, etc.
3. Stemming – Reducing words to their root form (e.g., playing → play).
4. Lemmatization – Converting words to dictionary base form (better →
good).
5. Named Entity Recognition (NER) – Identifying names, places, dates.
6. Sentiment Analysis – Finding emotions: positive, negative, neutral.
7. Parsing – Understanding sentence structure.
8. Machine Translation – Translating languages (e.g., English ↔ Hindi).
9. Text Summarization – Creating short summaries of long texts.
[Link] Recognition – Converting speech to text (e.g., Siri, Google
Assistant).

Applications of NLP
 Chatbots and Virtual Assistants (Alexa, Siri, Google Assistant)
 Social Media Monitoring (sentiment analysis, trend detection)
 Search Engines (Google understands queries)
 Spam Filtering
 Machine Translation (Google Translate)
 Customer Support Automation
 Medical text analysis (extracting symptoms, diagnosis)
 Document Summarization

Importance of NLP
 Helps automate text-heavy tasks
 Improves customer experience
 Extracts value from unstructured data
 Supports decision-making
 Enables human–computer interaction
 Useful for analyzing huge amounts of social media and review data

Text Mining Applications (Exam Answer)


Text mining is widely used across industries to extract valuable insights from
large amounts of unstructured text such as emails, social media posts, reviews,
documents, and reports. It helps organizations improve decision-making, detect
patterns, understand customer behavior, and identify trends.
1. Social Media Analytics
Text mining analyzes posts, comments, reviews, and hashtags on platforms like
Facebook, Instagram, YouTube, and Twitter. It helps identify trending topics,
understand public sentiment, and monitor brand reputation.

2. Customer Feedback & Review Analysis


Companies use text mining to analyze customer reviews from Amazon, Flipkart,
Google Reviews, and feedback forms. This helps identify customer satisfaction,
complaints, product issues, and areas for improvement.

3. Spam Detection & Email Filtering


Email systems use text mining to detect spam, phishing, and unwanted
messages. It classifies emails based on keywords, patterns, and writing style.

4. Fraud Detection
Banks and financial institutions use text mining to analyze suspicious messages,
claims, or transactions. It helps detect fraud patterns, insurance scams, and
abnormal behavior.

5. Healthcare & Medical Analysis


Text mining helps analyze clinical notes, medical reports, patient history, and
research papers. It supports disease diagnosis, treatment suggestions, and
identifying side effects.

6. Legal Document Analysis


Law firms use text mining to analyze contracts, case records, and legal
documents. It helps speed up research, identify relevant information, and
support case decisions.

7. Market Research & Competitive Analysis


Companies use text mining to analyze news articles, blogs, forums, and online
discussions. It helps understand market trends, competitor activities, and
consumer preferences.

8. Chatbots & Virtual Assistants


Text mining helps chatbots (like Siri, Alexa, customer support bots) understand
user queries and give accurate responses through NLP and intent recognition.

9. Topic Modeling & Trend Detection


Organizations use text mining to find hidden themes in documents and detect
emerging trends in industries such as finance, technology, and entertainment.

10. Document Classification


Used in libraries, digital repositories, and companies to automatically categorize
documents into predefined categories like finance, marketing, HR, etc.

Text Mining Process (Step-by-Step)


Text mining involves transforming unstructured text into meaningful insights.
The process includes several stages, from collecting text data to extracting
useful patterns.

1. Data Collection
The first step is gathering text data from various sources such as:
 Social media posts
 Customer reviews
 Emails
 Websites
 Documents
 Blogs and news articles
This becomes the raw input for text mining.

2. Text Preprocessing (Cleaning the Text)


Preprocessing removes noise to prepare text for analysis. It includes:
 Lowercasing text
 Removing stop words (the, is, a)
 Removing punctuation and numbers
 Handling emojis/special characters
 Removing extra spaces
This step improves the quality of the data.

3. Tokenization
Breaking text into smaller units called tokens (words or phrases).
Example: “Text mining is powerful” → [Text, mining, is, powerful]

4. Stemming and Lemmatization


These techniques reduce words to their base form.
 Stemming: playing → play
 Lemmatization: better → good
This helps standardize similar words.

5. Feature Extraction / Representation


Converts text into numerical form for machine learning. Common methods
include:
 Bag of Words (BoW)
 TF-IDF (Term Frequency–Inverse Document Frequency)
 Word Embeddings (Word2Vec, GloVe, BERT)
6. Text Mining Techniques / Modeling
Once text is converted into numerical features, various techniques are applied:
 Classification (spam or not spam)
 Clustering (grouping similar documents)
 Sentiment Analysis
 Topic Modeling (LDA)
 Information Extraction (names, places)
 Summarization

7. Evaluation and Validation


Models are checked for accuracy, performance, and reliability.
Metrics include:
 Precision
 Recall
 F1 score
 Accuracy

8. Visualization and Interpretation


Results are displayed using:
 Word clouds
 Charts and graphs
 Topic clusters
 Sentiment distribution
This helps decision-makers understand insights easily.
Text Mining Tools (Exam Answer)
Text mining tools are software applications and platforms that help extract
insights, patterns, and knowledge from large amounts of unstructured text data.
These tools support tasks like text preprocessing, NLP, sentiment analysis,
classification, clustering, topic modeling, and visualization.

1. RapidMiner
 A popular data mining and text mining tool.
 Provides drag-and-drop interface.
 Supports text preprocessing, sentiment analysis, classification, and
clustering.
 Easy to use for beginners.

2. Weka
 Open-source data mining software.
 Contains text mining plugins.
 Supports preprocessing, feature extraction, and machine learning models.
 Good for educational and research use.

3. KNIME (Konstanz Information Miner)


 Open-source analytics platform.
 Offers powerful text mining extensions.
 Supports tokenization, tagging, NER, topic modeling, and visualization.
 Easy integration with Python, R, and deep learning tools.

4. Python (NLTK, spaCy, TextBlob)


 Python libraries widely used in text mining and NLP.
 NLTK: Tokenization, stemming, parsing, POS tagging.
 spaCy: Fast NLP library for NER, dependency parsing, embeddings.
 TextBlob: Simple sentiment analysis and preprocessing.

5. R (tm, tidytext)
 R packages for text mining and visualization.
 tm: Text preprocessing, term-document matrices, stemming.
 tidytext: Modern text mining using tidy data principles.

6. SAS Text Miner


 Enterprise-level text mining software.
 Supports topic modeling, clustering, NER, and sentiment analysis.
 Often used in large organizations.

7. IBM Watson Natural Language Understanding


 Cloud-based AI tool for NLP and text analytics.
 Performs sentiment, emotion detection, entity extraction, and keyword
analysis.
 Highly accurate and scalable.

8. Google Cloud Natural Language API


 Provides sentiment analysis, entity detection, syntax parsing.
 Easy to integrate with applications.

9. LingPipe
 Java-based text mining tool.
 Used for NER, classification, and sentiment analysis.
 Popular in research and enterprise applications.
10. Orange Text Mining
 Visual programming tool.
 Supports text preprocessing, word clouds, clustering, topic modeling.

11. MATLAB Text Analytics Toolbox


 Provides tools for preprocessing, topic modeling, and sentiment analysis.

12. Apache OpenNLP


 Machine learning–based NLP framework.
 Supports tokenization, POS tagging, NER, parsing, and coreference
resolution.

You might also like