Text Analytics and Mining Explained
Text Analytics and Mining Explained
MODULE-5
Both terms deal with transforming unstructured text into structured, actionable
information using Natural Language Processing (NLP) and analytics techniques.
There are subtle differences:
Term Focus Scope
A broader concept involving the full Includes information retrieval,
Text
process of retrieving, extracting, and information extraction, data mining,
Analytics
analyzing text data and web mining
Text Mining
Text Mining (also known as Text Data Mining or Knowledge Discovery in Textual
Databases) is a semi-automated process that:
Extracts patterns and useful information from large volumes of unstructured text
Converts unstructured text into structured data for analysis
It is conceptually similar to Data Mining, but differs in the type of data processed:
Aspect Data Mining Text Mining
Structured (e.g., databases, Unstructured (e.g., documents, articles,
Data Type
spreadsheets) emails)
Input
Numeric/categorical values Text, XML, PDFs, Word files
Format
Impose structure on text and extract
Process Discover patterns in structured data
patterns
Classification, clustering, NLP, parsing, information extraction,
Techniques
association rules classification
Text mining has become crucial in areas where large text volumes are produced daily:
Domain Examples of Text Mining Use
Law Analyze court orders, judgments, and legal precedents
Academia Identify emerging research topics from scientific articles
Finance Analyze annual/quarterly financial reports
Medicine Extract data from patient discharge summaries or medical notes
Example:
Customer feedback and warranty claims can reveal product flaws and customer sentiment.
By analyzing such text, businesses can:
b. Topic Tracking
Predicts new documents that might interest a user based on their reading history or
profile.
Example: Recommending news articles similar to those already read.
c. Summarization
e. Clustering
f. Concept Linking
g. Question Answering
Technology Insights
Together, these fields form the foundation of text analytics and mining.
Real-World Impact
Application Case 5.1: Insurance Group Strengthens Risk Management with Text
Mining Solution
When asked for the biggest challenge facing the Czech automobile insurance industry, Peter
Jedlicˇka, PhD, doesn’t hesitate. “Bodily injury claims are growing disproportionately
compared with vehicle damage claims,” says Jedlicˇka, team leader of actuarial services for
the Czech Insurers’ Bureau (CIB). CIB is a professional organization of insurance companies
in the Czech Republic that handles uninsured, international, and untraced claims for what’s
known as motor third-party liability. “Bodily injury damages now represent about 45% of the
claims made against our members, and that proportion will continue to increase because of
recent legislative changes.” One of the difficulties that bodily injury claims pose for insurers
is that the extent of an injury is not always predictable in the immediate aftermath of a
vehicle accident. Injuries that were not at first obvious may become acute later, and
apparently minor injuries can turn into chronic conditions. The earlier that insurance
companies can accurately estimate their liability for medical damages, the more precisely
they can manage their risk and consolidate their resources. However, because the needed
information is contained in unstructured documents such as accident reports and witness
statements, it is extremely time consuming for individual employees to perform the needed
analysis. To expand and automate the analysis of unstructured accident reports, witness
statements, and claim narratives, CIB deployed a data analysis solution based on Dell
Statistica Data Miner and the Statistica Text Miner extension. Statistica Data Miner offers a
set of intuitive, user-friendly tools that are accessible even to nonanalysts. Application Case
5.1 Insurance Group Strengthens Risk Management with Text Mining Solution The solution
reads and writes data from virtually all standard file formats and offers strong, sophisticated
data cleaning tools. It also supports even novice users with query wizards, called Data Mining
Recipes, that help them arrive at the answers they need more quickly. With the Statistica Text
Miner extension, users have access to extraction and selection tools that can be used to index,
classify, and cluster information from large collections of unstructured text data, such as the
narratives of insurance claims. In addition to using the Statistica solution to make predictions
about future medical damage claims, CIB can also use it to find patterns that indicate
attempted fraud or to identify needed road safety improvements.
Improves Accuracy of Liability Estimates
Jedlicˇka expects the Statistica solution to greatly improve the ability of CIB to predict the
total medical claims that might arise from a given accident. “The Statistica solution’s data
mining and text mining capabilities are already helping us expose additional risk
characteristics, thus making it possible to predict serious medical claims in earlier stages of
the investigation,” he says. “With the Statistica solution, we can make much more accurate
estimates of total damages and plan accordingly.”
Expands Service Offerings to Members
Jedlicˇka is also pleased that the Statistica solution helps CIB offer additional services to its
member companies. “We are in a data-driven business,” he says. “With Statistica, we can
provide our members with detailed analyses of claims and market trends. Statistica also helps
us provide even stronger recommendations concerning claims reserves.”
Intuitive for Business Users
The intuitive Statistica tools are accessible by even nontechnical users. “The outputs of our
Statistica analyses are easy to understand for business users,” says Jedlicˇka. “Our business
users also find that the analysis results are in line with their own experience and
recommendations, so they readily see the value in the Statistica solution.”
Questions for Discussion
1. How can text analytics and mining be used to keep up with changing business needs of
insurance companies?
2. What were the challenges, the proposed solution, and the obtained results?
3. Can you think of other uses of text analytics and text mining for insurance companies?
Introduction
Natural Language Processing (NLP) is a crucial subfield of Artificial Intelligence (AI)
and Computational Linguistics.
It focuses on enabling computers to understand, interpret, and generate human language
in a way that is both meaningful and useful.
In the context of Text Mining and Text Analytics, NLP helps transform unstructured
textual data (words, sentences, paragraphs) into structured formats (numeric or symbolic
data) that can be easily analyzed by algorithms.
One of the earliest methods used in text mining was the Bag-of-Words (BoW) approach.
In this method:
Example:
Under the bag-of-words model, both sentences are represented by the same words {dogs,
chase, cats}, even though their meaning is different.
Even today, simple tasks such as spam filtering use the BoW model effectively.
Although this method works for basic categorization, it does not truly understand the
meaning or context of language.
c. Limitation of Bag-of-Words
For example, in medical text classification, researchers found that using BoW representation
on millions of abstracts from MEDLINE yielded poor results — nearly equivalent to random
guessing — because the model lacked contextual and semantic understanding.
Thus, researchers moved toward more advanced linguistic techniques, giving rise to NLP.
What is NLP?
Definition:
Natural Language Processing is a computational approach to analyzing and
understanding human language, converting textual information into structured data that
computers can process.
The goal of NLP is to achieve a “true understanding” of language — enabling machines not
just to process text, but to interpret its meaning.
Challenges in NLP
This greatly expanded the usability of WordNet while reducing manual effort.
Semantic search
Text classification
Sentiment analysis
Applications of NLP
Example:
NLP is used to analyze customer reviews, feedback, and social media posts.
It helps identify customer emotions toward products or services.
Businesses can detect dissatisfaction early, improve services, and target marketing
efforts effectively.
In simple terms:
That means NLP provides language structure and meaning, and data mining algorithms
discover patterns and insights from that structured data.
Application Case 5.2: AMC Networks Is Using Analytics to Capture New Viewers,
Predict Ratings, and Add Value for Advertisers in a Multichannel World
Over the past 10 years, the cable television sector in the United States has enjoyed a period of
growth that has enabled unprecedented creativity in the creation of high-quality content.
AMC Networks has been at the forefront of this new golden age of television, producing a
string of successful, critically acclaimed shows such as Breaking Bad, Mad Men, and The
Walking Dead. Dedicated to producing quality programming and movie content for more
than 30 years, AMC Networks Inc. owns and operates several of the most popular and award-
winning brands in cable television, producing and delivering distinctive, compelling, and
culturally relevant content that engages audiences across multiple platforms.
Getting Ahead of the Game
Despite its success, the company has no plans to rest on its laurels. As Vitaly Tsivin, SVP
Business Intelligence, explains: “We have no interest in standing still. Although a large
percentage of our business is still linear cable TV, we need to appeal to a new generation of
millennials who consume content in very different ways. “TV has evolved into a
multichannel, multistream business, and cable networks need to get smarter about how they
market to and connect with audiences across all of those streams. Relying on traditional
ratings data and third-party analytics providers is going to be a losing strategy: you need to
take ownership of your data, and use it to get a richer picture of who your viewers are, what
they want, and how you can keep their attention in an increasingly crowded entertainment
marketplace.”
Zoning in on the Viewer
The challenge is that there is just so much information available—hundreds of billions of
rows of data from industry data providers such as Nielsen and comScore, from channels such
as AMC’s TV Everywhere live Web streaming and video on-demand service, from retail
partners such as iTunes and Amazon, and from third-party online video services such as
Netflix and Hulu. “We can’t rely on high-level summaries; we need to be able to analyze
both structured and unstructured data, minute-by-minute and viewer-byviewer,” says Vitaly
Tsivin. “We need to know who’s watching and why—and we need to know it quickly so that
we can decide, for example, whether to run an ad or a promo in a particular slot during
tomorrow night’s episode of Mad Men.” AMC decided it needed to develop an
industryleading analytics capability in-house—and focused on delivering this capability as
quickly as possible. Instead of conducting a prolonged and expensive vendor and product
selection process, AMC decided to leverage its existing relationship with IBM as its trusted
strategic technology partner. The time and money traditionally spent on procurement were
instead invested in realizing the solution—accelerating AMC’s progress on its analytics
roadmap by at least 6 months.
Empowering the Research Department
In the past, AMC’s research team spent a large portion of its time processing data. Today,
thanks to its new analytics tools, it is able to focus most of its energy on gaining actionable
insights. “By investing in big data analytics technology from IBM, we’ve been able to
increase the pace and detail of our research an order of magnitude,” says Vitaly Tsivin.
“Analyses that used to take days and weeks are now possible in minutes, or even seconds.
“Bringing analytics in-house will provide major ongoing cost-savings. Instead of paying
hundreds of thousands of dollars to external vendors when we need some analysis done, we
can do it ourselves— more quickly, more accurately, and much more costeffectively. We’re
expecting to see a rapid return on investment. “As more sources of potential insight become
available and analytics becomes more strategic to the business, an in-house approach is really
the only viable way forward for any network that truly wants to gain competitive advantage
from its data.”
are also much more successful. In one recent example, intelligent segmentation and lookalike
modeling helped the company target new and existing viewers so effectively that AMC video
on-demand transactions were higher than would be expected otherwise. This newfound
ability to reach out to new viewers based on their individual needs and preferences is not just
valuable for AMC—it also has huge potential value for the company’s advertising partners.
AMC is currently working on providing access to its rich data sets and analytics tools as a
service for advertisers, helping them fine-tune their campaigns to appeal to ever-larger
audiences across both linear and digital channels. Vitaly Tsivin concludes: “Now that we can
really harness the value of big data, we can build a much more attractive proposition for both
consumers and advertisers—creating even better content, marketing it more effectively, and
helping it reach a wider audience by taking full advantage of our multichannel capabilities.”
Before this, both had separate databases; now, integration enables cross-agency data
analysis to detect threats efficiently.
iv. Deception Detection
Fuller, Biros, & Delen (2008) developed text mining models to detect deceptive vs.
truthful statements from criminal text data.
Achieved around 70% accuracy using only textual cues (no visual or vocal data).
Compared to polygraphs, this approach is non-intrusive and applicable to large text
corpora (like police interviews or online statements).
C. Biomedical Applications
The biomedical field produces vast amounts of literature and experimental data (e.g., gene,
protein, disease information). Text mining helps interpret and connect these findings.
i. Literature Analysis and Knowledge Discovery
Biomedical research generates enormous data through experiments like:
o DNA microarrays,
o SAGE (Serial Analysis of Gene Expression),
o Proteomics (mass spectrometry).
Scientists use text mining to integrate new experimental results with existing
biomedical literature for better understanding of biological entities.
ii. Protein Location Prediction
Shatkay et al. (2007) developed a text-mining system combining sequence-based
and text-based features to predict protein locations within cells.
Knowing where a protein resides helps understand its biological role and drug target
potential.
Their hybrid approach outperformed earlier models.
iii. Disease–Gene Relationship Extraction
Chun et al. (2006) built a system that extracts relationships between genes and
diseases from MEDLINE.
Used:
o A dictionary-based approach for gene/disease names.
o A machine learning–based Named Entity Recognition (NER) system to
filter false positives.
Result: 26.7% improvement in precision, with only a slight drop in recall.
iv. Gene–Protein Relationship Discovery
Nakov et al. (2005) illustrated how multi-level text analysis can uncover gene–
protein interactions.
The process involves:
1. Tokenization (breaking text into words),
2. Part-of-Speech tagging,
3. Parsing and matching terms to biological ontologies (structured domain
knowledge).
This helps decode complex relationships in projects like the Human Genome
Project.
D. Academic Applications
Text mining supports knowledge organization, indexing, and retrieval in academic and
publishing domains.
i. Scientific Publishing
Publishers maintain large digital databases needing automatic indexing for easier
retrieval.
Initiatives include:
o Nature’s Open Text Mining Interface – allows semantic querying.
o NIH Journal Publishing Document Type Definition – standardizes
document structures for automated mining.
ii. Academic Research Centers
National Centre for Text Mining (NaCTeM) – collaboration between the
University of Manchester and University of Liverpool.
o Initially focused on biomedical sciences, now extended to social sciences.
o Provides tools, training, and customized solutions for academic text mining.
BioText Project (UC Berkeley) – supports bioscience researchers in text analysis for
biomedical discovery.
iii. Importance in Academia
Helps researchers discover hidden connections, trends, and hypotheses.
Enables semantic search, where computers understand meaning, not just keywords.
credibility assessment) has involved face-to-face meetings and interviews. Yet, with the
growth of text-based communication, text-based deception-detection techniques are essential.
Techniques for successfully detecting deception—that is, lies—have wide applicability. Law
enforcement can use decision support tools and techniques to investigate crimes, conduct
security screening in airports, and monitor communications of suspected terrorists. Human
resources professionals might use deception-detection tools to screen applicants. These tools
and techniques also have the potential to screen e-mails to uncover fraud or other
wrongdoings committed by corporate officers. Although some people believe that they can
readily identify those who are not being truthful, a summary of deception research showed
that, on average, people are only 54% accurate in making veracity determinations (Bond &
DePaulo, 2006). This figure may actually be worse when humans try to detect deception in
text. Using a combination of text mining and data mining techniques, Fuller et al. (2008)
analyzed person-of-interest statements completed by people involved in crimes on military
bases. In these statements, suspects and witnesses are required to write their recollection of
the event in their own words. Military law enforcement personnel searched archival data for
statements that they could conclusively identify as being truthful or deceptive. These
decisions were made on the basis of corroborating evidence and case resolution. Once labeled
as truthful or deceptive, the law enforcement personnel removed identifying information and
gave the statements to the research team. In total, 371 usable statements were received for
analysis. The text-based deception-detection method used by Fuller et al. (2008) was based
on a process known as message feature mining, which relies on elements of data and text
mining techniques. A simplified depiction of the process is provided in Figure 5.3. First, the
researchers prepared the data for processing. The original handwritten statements had to be
transcribed into a word processing file. Second, features (i.e., cues) were identified. The
researchers identified 31 features representing categories or types of language that are
relatively independent of the text content and that can be readily analyzed by automated
means. For example, first-person pronouns such as I or me can be identified without analysis
of the surrounding text. Table 5.1 lists the categories and an example list of features used in
this study. The features were extracted from the textual statements and input into a flat file
for further processing. Using several feature-selection methods along with 10-fold cross-
validation, the researchers compared the prediction accuracy of three popular data mining
methods. Their results indicated that neural network models performed the best, with 73.46%
prediction accuracy on test data samples; decision trees performed second best, with 71.60%
accuracy; and logistic regression was last, with 65.28% accuracy. The results indicate that
automated text-based deception detection has the potential to aid those who must try to detect
lies in text and can be successfully applied to real-world data. The accuracy of these
techniques exceeded the accuracy of most other deception-detection techniques, even though
it was limited to textual cues.
Application Case 5.4: Bringing the Customer into the Quality Equation: Lenovo Uses
Analytics to Rethink Its Redesign
Lenovo was near final design on an update to the keyboard layout of one of its most popular
PCs when it spotted a small, but significant, online community of gamers who are
passionately supportive of the current keyboard design. Changing the design may have led to
a mass revolt of a large segment of Lenovo’s customer base—freelance developers and
gamers. The Corporate Analytics unit was using SAS as part of a perceptual quality project.
Crawling the web, sifting through text data for Lenovo mentions, the analysis unearthed a
previously unknown forum, where an existing customer had written a glowing six-page
review of the current design, especially the keyboard. The review attracted 2,000 comments!
“It wasn’t something we would have found in traditional preproduction design reviews,” says
Mohammed Chaara, Director of Customer Insight & VOC Analytics. It was the kind of
discovery that solidified Lenovo’s commitment to the Lenovo Early Detection (LED) system,
and the work of Chaara and his corporate analytics team. Lenovo, the largest global
manufacturer of PCs and tablets, didn’t set out to gauge sentiment around obscure bloggers or
discover new forums. The company wanted to inform quality, product development, and
product innovation by studying data—its own and that from outside the four walls. “We’re
mainly focused on supply chain optimization, crosssell/up-sell opportunities and pricing and
packaging of services. Any improvements we make in these areas are based on listening to
the customer,” Chaara says. SAS provides the framework to “manage the crazy amount of
data” that is generated. The project’s success has traveled like wildfire within the
organization. Lenovo initially planned on about 15 users, but word of mouth has led to 300
users signing up to log in to the LED dashboard for a visual presentation on customer
sentiment, warranty, and call center analysis.
The Results Have Been Impressive
• Over 50% reduction in issue detection time.
• 10 to 15% reduction in warranty costs from out-of-norm defects.
• 30 to 50% reduction in general information calls to the contact center.
Looking at the Big Picture
Traditional methods of gauging sentiment and understanding quality have built-in
weaknesses and time lags:
• Customer surveys only surface information from customers who are willing to fill them out.
• Warranty information often comes in months after delivery of the new product.
• It can be difficult to decipher myriad causes of customer discontent and product issues.
In addition, Lenovo sells its product packaged with software it doesn’t produce, and
customers use a variety of accessories (docking stations and mouse devices) that might or
might not be Lenovo products. To compound the issue, the company operates in 165
countries and supports more than 30 languages, so the manual methods to evaluate the
commentary were inconsistent, took too much time, and couldn’t scale to the volumes of
feedback it was seeing in social media. The sentiment analysis needed to be able to sense
nuances within the native languages. (For example, Australians describe things differently
than Americans.) The analysis-driven discovery of an issue with docking stations provided
the second big win for Lenovo’s LED initiative. Customers were calling tech support to say
they were having issues with the screen, or the machine shutting down abruptly, or the
battery wasn’t charging. Similar accounts were turning up on social media posts. Sometimes,
though not always, the customer mentioned docking. It wasn’t until Lenovo used SAS to
analyze the combination of call center notes and social media posts that the word docking
was connected to the problem, helping quality engineers figure out the root cause and issue a
software update. “We were able to pick up that feedback within weeks. It used to take 60 to
90 days because we had to wait for the reports to come back from the field,” Chaara says.
Now it takes just 15 to 30 days. That reduction in detection time has driven a 10 to 15%
reduction in warranty costs for those issues. As warranty claims cost the company about $1.2
billion yearly, this is a significant savings. Although the call center information was crucial,
the social media component was what sealed the deal. “With Twitter and Facebook, people
described what they were doing at that minute, “I docked the machine and X happened.’ It’s
raw, unbiased and so powerful,” Chaara says. An unforeseen insight was found when
analyzing what customers were saying as they got their PCs up and running. Lenovo realized
its documentation to explain its products, warranties, and the like was unclear. “There is a
cost to every call center call. With the improved documentation, we’ve seen a 30 to 50%
reduction in calls coming in for general information,” Chaara said.
Winning Praise beyond the frontlines
The project has been so successful that Chaara demoed it for the CEO. The goal is to
configure a dashboard view for the C-suite. “That’s the level of thinking from our senior
executives. They believe in this,” Chaara says. In addition, Chaara’s group will be formally
measuring the success of the effort and expanding it to measure issues like customer
experience when buying a Lenovo product. “The application of analytics has ultimately led
us to a more holistic understanding of the concept of quality. Quality isn’t just a PC working
correctly. It’s people knowing how to use it, getting quick and accurate help from the
company, getting the non-Lenovo components to work well with the hardware, and
understanding what the customers like about the existing product—rather than just
redesigning it because product designers think it’s the right thing to do. “SAS has allowed us
to get a definition of quality from the view of the customer,” Chaara says.
Questions for Discussion
1. How did Lenovo use text analytics and text mining to improve quality and design of their
products and ultimately improve customer satisfaction?
2. What were the challenges, the proposed solution, and the obtained results?
Commercial text mining software can import such data and convert it into flat files for
analysis.
Output:
A well-organized and standardized collection of text documents (the corpus) ready for
processing.
[Link]
Definition:
Groups similar documents together into clusters without predefined labels.
Benefits:
Type Explanation
Improved Recall Retrieves all relevant documents by considering similar ones.
Improved Precision Returns only the most relevant clusters to the user.
Popular Methods:
Scatter/Gather Clustering: Dynamically clusters documents for browsing when no
specific query is given.
Query-Specific Clustering: Hierarchically clusters documents related to a search
query, showing different levels of relevance.
3. Association
Definition:
Discovers relationships or co-occurrences among concepts or terms in documents.
Measures Used:
Support: % of documents where terms appear together.
Confidence: Likelihood that one term appears given another.
Example:
If “Software Implementation Failure” frequently appears with “ERP” and “CRM,”
then these terms have a strong association.
Applications:
Literature analysis (e.g., tracking topics like “bird flu” and related terms such as
“virus,” “vaccine,” “countries affected”).
Marketing: finding terms that co-occur in customer feedback.
4. Trend Analysis
Definition:
Analyzes how the frequency and relationships of terms evolve over time.
Example:
Tracking how concepts like “AI,” “cloud computing,” or “data privacy” have evolved in
academic journals over years.
Process:
Compare term distributions across time-based collections.
Identify emerging or declining topics.
Case Example:
Delen & Crossland (2008) used trend analysis on top information systems journals to
observe how research themes evolved.
be even more tedious and complex. Trying to ferret out relevant work that others have
reported may be difficult, at best, and perhaps even near impossible if traditional, largely
manual reviews of published literature are required. Even with a legion of dedicated graduate
students or helpful colleagues, trying to cover all potentially relevant published work is
problematic. Many scholarly conferences take place every year. In addition to extending the
body of knowledge of the current focus of a conference, organizers often desire to offer
additional minitracks and workshops. In many cases, these additional events are intended to
introduce the attendees to significant streams of research in related fields of study and to try
to identify the “next big thing” in terms of research interests and focus. Identifying
reasonable candidate topics for such minitracks and workshops is often subjective rather than
derived objectively from the existing and emerging research. In a recent study, Delen and
Crossland (2008) proposed a method to greatly assist and enhance the efforts of the
researchers by enabling a semiautomated analysis of large volumes of published literature
through the application of text mining. Using standard digital libraries and online publication
search engines, the authors downloaded and collected all the available articles for the three
major journals in the field of management information systems: MIS Quarterly (MISQ),
Information Systems Research (ISR), and the Journal of Management Information Systems
(JMIS). To maintain the same time interval for all three journals (for potential comparative
longitudinal studies), the journal with the most recent starting date for its digital publication
availability was used as the start time for this study (i.e., JMIS articles have been digitally
available since 1994). For each article, they extracted the title, abstract, author list, published
keywords, volume, issue number, and year of publication. They then loaded all the article
data into a simple database file. Also included in the combined data set was a field that
designated the journal type of each article for likely discriminatory analysis. Editorial notes,
research notes, and executive overviews were omitted from the collection. Table 5.2 shows
how the data was presented in a tabular format. In the analysis phase, they chose to use only
the abstract of an article as the source of information extraction. They chose not to include
the keywords listed with the publications for two main reasons: (1) under normal
circumstances, the abstract would already include the listed keywords, and therefore
inclusion of the listed keywords for the analysis would mean repeating the same information
and potentially giving them unmerited weight; and (2) the listed keywords may be terms that
authors would like their article to be associated with (as opposed to what is really contained
in the article), therefore potentially introducing unquantifiable bias to the analysis of the
content. The first exploratory study was to look at the longitudinal perspective of the three
journals (i.e., evolution of research topics over time). In order to conduct a longitudinal study,
they divided the 12-year period (from 1994 to 2005) into four 3-year periods for each of the
three journals. This framework led to 12 text mining experiments with 12 mutually exclusive
data sets. At this point, for each of the 12 data sets they used text mining to extract the most
descriptive terms from these collections of articles represented by their abstracts. The results
were tabulated and examined for time-varying changes in the terms published in these three
journals. As a second exploration, using the complete data set (including all three journals
and all four periods), they conducted a clustering analysis. Clustering is arguably the most
commonly used text mining technique. Clustering was used in this study to identify the
natural groupings of the articles (by putting them into separate clusters) and then to list the
most descriptive terms that characterized those clusters. They used SVD to reduce the
dimensionality of the term-by-document matrix and then an expectation-maximization
algorithm to create the clusters. They conducted several experiments to identify the optimal
number of clusters, which turned out to be nine. After the construction of the nine clusters,
they analyzed the content of those clusters from two perspectives: (1) representation of the
journal type (see Figure 5.8a) and (2) representation of time (Figure 5.8b). The idea was to
explore the potential differences and/or commonalities among the three journals and potential
changes in the emphasis on those clusters; that is, to answer questions such as “Are there
clusters that represent different research themes specific to a single journal?” and “Is there a
time-varying characterization of those clusters?” They discovered and discussed several
interesting patterns using tabular and graphical representation of their findings (for further
information see Delen & Crossland, 2008).
Table: Tabular Representation of the Fields Included in the Combined Data Set
FIGURE: (a) Distribution of the Number of Articles for the Three Journals over the Nine
Clusters; (b) Development of the Nine Clusters over the Years.
Sentiment Analysis
Introduction to Sentiment Analysis
Sentiment Analysis (SA) — also known as Opinion Mining or Subjectivity Analysis — is
a text analytics technique used to identify, extract, and quantify subjective information
such as opinions, emotions, and attitudes expressed in textual data.
In simple terms, sentiment analysis answers the question:
Dept of AI&DS, SIET Page 26
BUSINESS ANALYTICS BAD714B
“What do people feel or think about a certain topic, product, person, or event?”
It transforms unstructured text (like tweets, reviews, posts, and blogs) into structured data
that reflects public sentiment — positive, negative, or neutral.
4. Brand Management
Monitors online platforms to track brand reputation.
Detects negative publicity early and manages crises.
Companies use social listening tools to monitor mentions and opinions.
5. Financial Markets
Investor sentiment strongly influences stock prices.
SA helps predict short-term market movements based on:
o News headlines
o Social media buzz
o Discussion forums
Example: Detecting market panic or optimism before trading decisions.
6. Politics
Analyzes public opinions during elections or debates.
Predicts voter preferences, popularity of candidates, and reaction to policies.
Used successfully in U.S. presidential campaigns (2008, 2012).
7. Government Intelligence
Monitors public reactions to policies, proposals, or events.
Detects hostile or negative sentiment spikes for security or regulatory alerts.
Useful for agencies like Homeland Security.
Application Case 5.6: Creating a Unique Digital Experience to Capture the Moments
That Matter at Wimbledon
channels. The constant evolution of these digital platforms is the result of a 26-year
partnership between the AELTC and IBM. Mick Desmond, Commercial and Media Director
at the AELTC, explains: “When you watch Wimbledon on TV, you are seeing it through the
broadcaster’s lens. We do everything we can to help our media partners put on the best
possible show, but at the end of the day, their broadcast is their presentation of The
Championships. “Digital is different: it’s our platform, where we can speak directly to our
fans—so it’s vital that we give them the best possible experience. No sporting event or media
channel has the right to demand a viewer’s attention, so if we want to strengthen our brand,
we need people to see our digital experience as the number-one place to follow The
Championships online.” To that end, the AELTC set a target of attracting 70 million visits, 20
million unique devices, and 8 million social followers during the two weeks of The
Championships 2015. It was up to IBM and AELTC to find a way to deliver.
Delivering a Unique Digital Experience
IBM and the AELTC embarked on a complete redesign of the digital platform, using their
intimate knowledge of The Championships’ audience to develop an experience tailor-made to
attract and retain tennis fans from across the globe. “We recognized that while mobile is
increasingly important, 80% of our visitors are using desktop computers to access our
website,” says Alexandra Willis, Head of Digital and Content at the AELTC. “Our challenge
for 2015 was how to update our digital properties to adapt to a mobilefirst world, while still
offering the best possible desktop experience. We wanted our new site to take maximum
advantage of that large screensize and give desktop users the richest possible experience in
terms of high-definition visuals and video content—while also reacting and adapting
seamlessly to smaller tablet or mobile formats. “Second, we placed a major emphasis on
putting content in context—integrating articles with relevant photos, videos, stats and
snippets of information, and simplifying the navigation so that users could move seamlessly
to the content that interests them most.” On the mobile side, the team recognized that the
wider availability of high-bandwidth 4G connections meant that the mobile Website would
become more popular than ever—and ensured that it would offer easy access to all rich media
content. At the same time, The Championships’ mobile apps were enhanced with real-time
notifications of match scores and events—and could even greet visitors as they passed
through stations on the way to the grounds. The team also built a special set of Websites for
the most important tennis fans of all: the players themselves. Using IBM® Bluemix®
technology, it built a secure Web application that provided players with a personalized view
of their court bookings, transport, and on-court times, as well as helping them review their
performance with access to stats on every match they played.
Turning Data into Insight—and Insight into Narrative
To supply its digital platforms with the most compelling possible content, the team took
advantage of a unique advantage: its access to real-time, shotby-shot data on every match
played during The Championships. Over the course of the Wimbledon fortnight, 48 courtside
experts capture approximately 3.4 million datapoints, tracking the type of shot, the strategies,
and the outcome of each and every point. This data is collected and analyzed in real time to
produce statistics for TV commentators and journalists—and also for the digital platform’s
own editorial team. “This year IBM gave us an advantage that we had never had before—
using data streaming technology to provide our editorial team with real-time insight into
significant milestones and breaking news,” says Alexandra Willis. “The system automatically
watched the streams of data coming in from all 19 courts, and whenever something
significant happened—such as Sam Groth hitting the second-fastest serve in Championships,
history—it let us know instantly. Within seconds, we were able to bring that news to our
digital audience and share it on social media to drive even more traffic to our site. “The
ability to capture the moments that matter and uncover the compelling narratives within the
data, faster than anyone else, was key. If you wanted to experience the emotions of The
Championships live, the next best thing to being there in person was to follow the action on
[Link].”
Harnessing the Power of Natural Language
Another new capability trialed this year was the use of IBM’s NLP technologies to help mine
the AELTC’s huge library of tennis history for interesting contextual information. The team
trained IBM Watson™ Engagement Advisor to digest this rich unstructured data set and use
it to answer queries from the press desk. The same NLP front-end was also connected to a
comprehensive structured database of match statistics, dating back to the first Championships
in 1877—providing a one-stop shop for both basic questions and more complex inquiries.
“The Watson trial showed a huge amount of potential. Next year, as part of our annual
innovation planning process, we will look at how we can use it more widely—ultimately in
pursuit of giving fans more access to this incredibly rich source of tennis knowledge,” says
Mick Desmond.
Taking to the Cloud
The whole digital environment was hosted by IBM in its Hybrid Cloud. IBM used
sophisticated modeling techniques to predict peaks in demand based on the schedule, the
popularity of each player, the time of day, and many other factors—enabling it to
dynamically allocate cloud resources appropriately to each piece of digital content and ensure
a seamless experience for millions of visitors around the world. In addition to the powerful
private cloud platform that has supported The Championships for several years, IBM also
used a separate SoftLayer® cloud to host the Wimbledon Social Command Centre and also
provide additional incremental capacity to supplement the main cloud environment during
times of peak demand. The elasticity of the cloud environment is key, as The Championships’
digital platforms need to be able to scale efficiently by a factor of more than 100 within a
matter of days as the interest builds ahead of the first match on Centre Court.
Keeping Wimbledon Safe and Secure
Online security is a key concern nowadays for all organizations. For major sporting events in
particular, brand reputation is everything—and while the world is watching, it is particularly
important to avoid becoming a high-profile victim of cyber-crime. For these reasons, security
has a vital role to play in IBM’s partnership with the AELTC. Over the first five months of
2015, IBM security systems detected a 94% increase in security events on the
[Link] infrastructure, compared to the same period in 2014. As security threats—
and in particular distributed denial of service (DDoS) attacks—become ever more prevalent,
IBM continually increases its focus on providing industry-leading levels of security for the
AELTC’s whole digital platform. A full suite of IBM security products, including IBM
QRadar® SIEM and IBM Preventia Intrusion Prevention, enabled this year’s Championships
to run smoothly and securely and the digital platform to deliver a high-quality user
experience at all times.
Capturing Hearts and Minds
The success of the new digital platform for 2015— supported by IBM cloud, analytics,
mobile, social, and security technologies—was immediate and complete. Targets for total
visits and unique visitors were not only met, but exceeded. Achieving 71 million visits and
542 million page views from 21.1 million unique devices demonstrates the platform’s success
in attracting a larger audience than ever before and keeping those viewers engaged
throughout The Championships. “Overall, we had 13% more visits from 23% more devices
than in 2014, and the growth in the use of [Link] on mobile was even more
impressive,” says Alexandra Willis. “We saw 125% growth in unique devices on mobile,
98% growth in total visits, and 79% growth in total page views.” Mick Desmond concludes:
“The results show that in 2015, we won the battle for fans’ hearts and minds. People may
have favorite newspapers and sports websites that they visit for 50 weeks of the year—but for
two weeks, they came to us instead. “That’s a testament to the sheer quality of the experience
we can provide—harnessing our unique advantages to bring them closer to the action than
any other media channel. The ability to capture and communicate relevant content in real
time helped our fans experience The Championships more vividly than ever before.”
Questions for Discussion
1. How did Wimbledon use analytics capabilities to enhance viewers’ experience?
2. What were the challenges, the proposed solution, and the obtained results?
Topic Modeling
Definition:
Topic modeling is an advanced text mining technique used to automatically discover
hidden thematic structures (topics) from a large collection of unstructured text documents.
It identifies groups of words that frequently occur together and represent a particular concept
or theme within the corpus.
Purpose of Topic Modeling
To organize and summarize large collections of textual data.
To uncover abstract “topics” discussed across multiple documents.
To help in document classification, summarization, and recommendation systems.
To support decision-making by identifying key themes, trends, or patterns in text data.
How It Works
Topic modeling assumes that:
Each document is composed of multiple topics.
Each topic is characterized by a distribution of words.
Example:
Document: “The movie had great visual effects but poor acting.”
Topics identified:
• Entertainment (words: movie, acting, performance)
• Technology (words: visual, effects, graphics)
Data mining
Text mining
Information retrieval
Machine learning
Natural language processing (NLP)
Researchers use web crawlers to automatically collect data about thousands of movies —
including box office revenue, genre, cast, and ratings — to predict financial success using
data mining models.
Fraud detection
Predicting customer churn or conversion
Typical Techniques
Clustering: Group users with similar browsing patterns.
Association rules: Discover frequently visited page sequences.
Sequential pattern mining: Identify navigation paths.
Advantages
Extracts actionable knowledge from the vast Web.
Enables real-time business decision making.
Enhances personalization and customer experience.
Improves competitiveness through data-driven insights.
Limitations
Privacy and ethical concerns (tracking user data).
Dynamic nature of Web content leads to data inconsistency.
High computational cost and storage requirements.
Data quality issues (spam, redundancy, or missing metadata).
Search Engine
A search engine is a software system that searches information stored on the Web and
returns results based on the user’s query.
Examples: Google, Bing, Yahoo, DuckDuckGo.
When users type keywords or a sentence, the search engine returns Web pages that are most
relevant to those keywords.
Goals of a Search Engine
Goal Meaning
Effectiveness (Quality) Return the most relevant and useful search results.
Efficiency (Speed) Return results quickly.
These two goals are balanced to provide the best user experience.
A. Development Cycle
It includes:
1. Web Crawler (Spider)
2. Document Indexer
i. Web Crawler
Automatically browses and downloads Web pages.
Starts from a list of URLs (called seeds).
Finds new pages by following hyperlinks.
Stores these pages for processing.
B. Response Cycle
It includes:
1. Query Analyzer
2. Document Matcher/Ranker
i. Query Analyzer
Takes user’s search text.
Converts it to same structure used in indexing.
Performs:
o Tokenization
o Stop-word removal
o Stemming
o Spell correction / synonym checks
ii. Document Matcher and Ranker
Matches processed query against indexed database.
Ranks documents by relevance.
Page Ranking
Originally, search engines only matched keywords — quality was poor.
Google introduced PageRank Algorithm (1997):
A page is important if many other important pages link to it.
Similar to academic citation: highly cited pages are more influential.
This ranking, combined with content relevance, improves search result quality.
SEO Techniques
Technique Purpose
Cross-linking pages internally Strengthens important pages
Use relevant keywords in content Helps match user queries
Updating content regularly Encourages frequent crawling
Adding keywords in title and meta descriptions Helps search engine relevance
Increasing backlinks from other sites Improves authority and PageRank
Application Case 5.7: Understanding Why Customers Abandon Shopping Carts Results
in a $10 Million Sales Increase
[Link], the leading Internet shopping mall in Korea with 13 million customers, has
developed an integrated Web traffic analysis system using SAS for Customer Experience
Analytics. As a result, Lotte. com has been able to improve the online experience for its
customers, as well as generate better returns from its marketing campaigns. Now, Lotte. com
executives can confirm results anywhere, anytime, as well as make immediate changes. With
almost one million Web site visitors each day, [Link] needed to know how many visitors
were making purchases and which channels were bringing the most valuable traffic. After
reviewing many diverse solutions and approaches, [Link] introduced its integrated Web
traffic analysis system using the SAS for Customer Experience Analytics solution. This is the
first online behavioral analysis system applied in Korea. With this system, [Link] can
accurately measure and analyze Web site visitor numbers, page view status of site visitors
and purchasers, the popularity of each product category and product, clicking preferences for
each page, the effectiveness of campaigns, and much more. This information enables
[Link] to better understand customers and their behavior online, and conduct
sophisticated, costeffective targeted marketing. Commenting on the system, Assistant
General Manager Jung Hyo-hoon of the Marketing Planning Team for [Link] said, “As a
result of introducing the SAS system of analysis, many ‘new truths’ were uncovered around
customer behavior, and some of them were ‘inconvenient truths.’” He added, “Some site-
planning activities that had been undertaken with the expectation of certain results actually
had a low reaction from customers, and the site planners had a difficult time recognizing
these results.”
Benefits
Introducing the SAS for Customer Experience Analytics solution fully transformed the
[Link] Web site. As a result, [Link] has been able to improve the online experience
for its customers as well as generate better returns from its marketing campaigns. Since
implementing SAS for Customer Experience Analytics, [Link] has seen many benefits.
A Jump in Customer Loyalty
A large amount of sophisticated activity information can be collected under a visitor
environment, including quality of traffic. Deputy Assistant General Manager Jung said that
“by analyzing actual valid traffic and looking only at one to two pages, we can carry out
campaigns to heighten the level of loyalty, and determine a certain range of effect,
accordingly.” He added, “In addition, it is possible to classify and confirm the order rate for
each channel and see which channels have the most visitors.”
Optimized Marketing Efficiency Analysis
Rather than just analyzing visitor numbers only, the system is capable of analyzing the
conversion rate (shopping cart, immediate purchase, wish list, purchase completion)
compared to actual visitors for each campaign type (affiliation or e-mail, banner, keywords,
and others), so detailed analysis of channel effectiveness is possible. In addition, it can
confirm the most popular search words used by visitors for each campaign type, location, and
purchased products. The page overlay function can measure the number of clicks and number
of visitors for each item in a page to measure the value for each location in a page. This
capability enables [Link] to promptly replace or renew low-traffic items.
Enhanced Customer Satisfaction and Customer Experience Lead to Higher Sales
[Link] built a customer behavior analysis database that measures each visitor, what pages
are visited, how visitors navigate the site, and what activities are undertaken to enable diverse
analysis and improve site efficiency. In addition, the database captures customer
demographic information, shopping cart size and conversion rate, number of orders, and
number of attempts. By analyzing which stage of the ordering process deters the most
customers and fixing those stages, conversion rates can be increased. Previously, analysis
was done only on placed orders. By analyzing the movement pattern of visitors before
ordering and at the point where breakaway occurs, customer behavior can be forecast, and
sophisticated marketing activities can be undertaken. Through a pattern analysis of visitors,
purchases can be more effectively influenced and customer demand can be reflected in real
time to ensure quicker responses. Customer satisfaction has also improved as Lotte. com has
better insight into each customer’s behaviors, needs, and interests. Evaluating the system,
Jung commented, “By finding out how each customer group moves on the basis of the data, it
is possible to determine customer service improvements and target marketing subjects, and
this has aided the success of a number of campaigns.” However, the most significant benefit
of the system is gaining insight into individual customers and various customer groups. By
understanding when customers will make purchases and the manner in which they navigate
throughout the Web page, targeted channel marketing and better customer experience can
now be achieved. Plus, when SAS for Customer Experience Analytics was implemented by
[Link]’s largest overseas distributor, it resulted in a first-year sales increase of 8 million
euros (US$10 million) by identifying the causes of shopping cart abandonment.
Questions for Discussion
1. How did [Link] use analytics to improve sales?
2. What were the challenges, the proposed solution, and the obtained results?
Social Analytics
Social Analytics refers to the monitoring, analyzing, measuring, and interpreting digital
interactions among people on social platforms. Its purpose is to extract business insights
from social behavior and communication patterns.
It focuses on:
What people say (text, comments, posts, reviews)
How people interact (likes, shares, follows, influence)
How groups form and evolve online
According to Gartner, social analytics includes:
Measuring conversations
Tracking relationships
Understanding community influence
Predicting behaviors and trends
SNA Metrics
SNA uses several quantitative metrics to analyze networks:
A. Connections-Based Metrics
Metric Meaning
Homophily People tend to connect with similar individuals
Strength increases when ties have multiple roles (e.g.,
Multiplexity
coworker + friend)
Reciprocity Mutual relationships (both follow each other)
Network Closure /
Your friends are also friends with each other
Transitivity
People connect with geographically or socially close
Propinquity
individuals
B. Distribution-Based Metrics
Metric Meaning Use
Centrality (Degree, Closeness, Who is most important in the
Find influencers
Betweenness) network
A person connecting two Useful for spreading
Bridge
groups messages
How interconnected the
Density Shows cohesion
network is
Strong vs. weak relationship Weak ties spread new
Tie Strength
bonds ideas faster
C. Segmentation-Based Metrics
Metric Meaning
Cliques Groups where everyone is connected to everyone
Clustering Coefficient Likelihood of group formation
Cohesion Minimum people needed to break the network
Application Case 5.8: Tito’s Vodka Establishes Brand Loyalty with an Authentic Social
Strategy
If Tito’s Handmade Vodka had to identify a single social media metric that most accurately
reflects its mission, it would be engagement. Connecting with vodka lovers in an inclusive,
authentic way is something Tito’s takes very seriously, and the brand’s social strategy reflects
that vision. Founded nearly two decades ago, the brand credits the advent of social media
with playing an integral role in engaging fans and raising brand awareness. In an interview
with Entrepreneur, founder Bert “Tito” Beveridge credited social media for enabling Tito’s to
compete for shelf space with more established liquor brands. “Social media is a great
platform for a word-of-mouth brand, because it’s not just about who has the biggest
megaphone,” Beveridge told Entrepreneur. As Tito’s has matured, the social team has
remained true to the brand’s founding values and actively uses Twitter and Instagram to have
one-onone conversations and connect with brand enthusiasts. “We never viewed social media
as another way to advertise,” said Katy Gelhausen, Web & Social Media Coordinator. “We’re
on social so our customers can talk to us.” To that end, Tito’s uses Sprout Social to
understand the industry atmosphere, develop a consistent social brand, and create a dialogue
with its audience. Recently and as a result, Tito’s organically grew its Twitter and Instagram
communities by 43.5% and 12.6%, respectively, within 4 months.
Informing a Seasonal, Integrated Marketing Strategy
Tito’s quarterly cocktail program is a key part of the brand’s integrated marketing strategy.
Each quarter, a cocktail recipe is developed and distributed through Tito’s online and offline
marketing initiatives. It is important for Tito’s to ensure the recipe is aligned with the brand’s
focus as well as larger industry direction. Therefore, Gelhausen uses Sprout’s Brand
Keywords to monitor industry trends and cocktail flavor profiles. “Sprout has been a really
important tool for social monitoring. The Inbox is a nice way to keep on top of hashtags and
see general trends in one stream,” said Gelhausen. These learnings are presented to Tito’s in-
house mixology team and used to ensure the same quarterly recipe is communicated to the
brand’s sales team and across marketing channels. “Whether you’re drinking Tito’s at a bar,
buying it from a liquor store or following us on social media you’re getting the same
quarterly cocktail,” said Gelhausen. The program ensures that, at every consumer touchpoint,
a person is receiving a consistent brand experience—and that consistency is vital. In fact,
according to an Infosys study on the omnichannel shopping experience, 34% of consumers
attribute cross-channel consistency as a reason they spend more with a brand. Meanwhile,
39% cite inconsistency as a reason enough to spend less. At Tito’s, gathering industry
insights starts with social monitoring on Twitter and Instagram through Sprout. But the
brand’s social strategy doesn’t stop there. Staying true to its roots, Tito’s uses the platform on
a daily basis to authentically connect with customers. Sprout’s Smart Inbox displays Tito’s
Twitter and Instagram accounts in a single, cohesive feed. This helps Gelhausen manage
inbound messages and quickly identify which require a response. “Sprout allows us to stay on
top of the conversations we’re having with our followers. I love how you can easily interact
with content from multiple accounts in one place,” said Gelhausen.
Spreading the Word on Twitter
Tito’s approach to Twitter is simple: engage in personal, one-on-one conversations with fans.
Dialogue is a driving force for the brand, and over the course of 4 months, 88% of Tweets
sent were replies to inbound messages. Using Twitter as an open line of communication
between Tito’s and its fans resulted in a 162.2% increase in engagement and a 43.5% gain in
followers. Even more impressively, Tito’s ended the quarter with 538,306 organic
impressions—an 81% rise. A similar strategy is applied to Instagram, which Tito’s uses to
strengthen and foster a relationship with fans by publishing photos and videos of new recipe
ideas, brand events and initiatives.
Capturing the Party on Instagram
On Instagram, Tito’s primarily publishes lifestyle content and encourages followers to
incorporate the brand into everyday occasions. Tito’s also uses the platform to promote its
cause marketing efforts and to tell its brand story. The team finds value in Sprout’s Instagram
Profiles Report, which helps them identify what media is receiving the most engagement,
analyze audience demographics and growth, dive deeper into publishing patterns, and
quantify outbound hashtag performance. “Given Instagram’s new personalized feed, it’s
important that we pay attention to what really does resonate,” said Gelhausen. Using the
Instagram Profiles Report, Tito’s has been able to measure the impact of its Instagram
marketing strategy and revise its approach accordingly. By utilizing the network as another
way to engage with fans, the brand has steadily grown its organic audience. In 4 months,
@TitosVodka saw a 12.6% rise in followers and a 37.1% increase in engagement. On
average, each piece of published content gained 534 interactions, and mentions of the brand’s
hashtag, #titoshandmadevodka, grew by 33%.
Where to from Here?
Social is an ongoing investment in time and attention. Tito’s will continue the momentum the
brand experienced by segmenting each quarter into its own campaign. “We’re always getting
smarter with our social strategies and making sure that what we’re posting is relevant and
resonates,” said Gelhausen. Using social to connect with fans in a consistent, genuine, and
memorable way will remain a cornerstone of the brand’s digital marketing efforts. Using
Sprout’s suite of social media management tools, Tito’s will continue to foster a community
of loyalists. Highlights:
• A 162% increase in organic engagement on Twitter
• An 81% increase in organic Twitter impressions
• A 37% increase in engagement on Instagram