Chapter 4
Social Media & Text analytics
Contents
• Overview of Social media analytics
• Key social media analytics methods
• Social network analysis
• Introduction to NLP
• Text analytics
• Trend analytics
• Challenges in Social media analytics
Overview of Social media analytics
• Social media apps/websites allow users to
– Interact in various ways to access information
– Allow users to contribute to information
• Social media sites generate and store a huge
amount of information, in various formats.
• Provides content-rich information for analysis.
• Social media analytics, one of the emerging
topics in the area of research in Data science.
• Social media analytics process consists of
three stages :
– Data capturing
– Data understanding
– Data presentation.
• Data capturing stage
– Gathering data from multiple sources
– Preprocessing the gathered data
– Extracting relevant information from the
preprocessed data
• Data understanding
– Removing noisy data
– Performing required analytics (like sentiment
analysis, recommender analysis etc)
• Data Presentation
– Summarize and generate the findings
– Present the findings, in a visualization manner.
Seven layers of social media analytics
• Social media consists of 7 layers of data , that
contain useful information that is often
garnered for business intelligence.
• These 7 layers of data may be
– Either visible( textual data) OR
– Invisible ( hyperlinks network)
• All 7 layers contribute substantially to social
media input, for gaining useful insights
• 7 layers
– Text
– Networks
– Actions
– Hyperlinks
– Mobile
– Location
– Search engine
• Text Layer 1 Textual messages includes
tweets, textual posts, comments, status updates
etc.
– Useful in business analytics to identify user
opinion/sentiment regarding a particular product,
topic, individual etc.
• Networks Layer 2 Focus on the networking
structure of the social media data.
– Indicates connection between users based on concept
of friends, followers etc.
– Often found in networking sites like Facebook, Insta
etc
– Network analysis done to identify influential users,
predict new links etc
• Actions Layer 3 Include actions performed by
users while using social meida, like button clicks,
post sharing , creating new groups/events, accept
friend requests etc
– Analysis done to find popularity of a person/item,
recent trends followed by users, popularity of user
groups etc.
• Mobile Layer 4 focus on the analysis of user
engagement with mobile applications.
– Mobile analytics usually done for marketing analysis
to attract those users who are highly engaged with a
mobile application
– In-app analysis concentrates on the kind of activities
and interaction of users with an app.
• Hyperlinks Layer 5 Commonly found in almost all
websites, that allow navigation from one page to
another.
– In-Link Hyperlink into a page
– Out-Link Hyperlink out of a web page.
– Hyperlink analytics is all about analyzing ad interpreting
social media hyperlinks
• Location Layer 6 Also known as geospatial
analytics.
– Done to gain insight from the geographical content of the
social media data.
– Eg Real time location analysis done by courier services used
by social media sites to keep track of the locations of
delivery in real-time
– Historical location analysis done to increase sales and profit
• Search Engines Layer 7 Includes analyzing
historial search data to generate informative
search statistics.
– Search statistics used for Search engine
optimization (SEO) and Search engine marketing
(SEM).
– Eg : Advertisement spending statistics, keyword
monitoring, trends analysis etc.
Social Media analytics lifecycle
• Social media analytics cycle consists of 6 steps
namely :
– Identification, extraction, cleaning, analyzing,
visualization and Interpretation.
• The core business objectives are displayed at
the centre, and it notifies each step of the
analytics cycle.
• 1. Identification identify the correct source
of data for doing analysis
• 2. Extraction Once the data source is
identified , the next step is to use a suitable
API for data extraction.
• 3. Cleaning Cleaning done as a data
preprocessing step to remove unwanted data ,
thus reducing the quantity of data to be
analyzed.
• 4. Analyzing Clean data is analyzed for
generating meaningful business insights.
• Visualization Presentation of the analyzed
data in a graphical format.
– Human brains process visual content better than
plain text.
– Interaction data visualization helps in easy
analysis and decision making.
• Interpretation
– Interpretation of the results
– Involves translating the outcomes into meaningful
business solutions.
Key Social media analytics methods
• 3 primary methods for social media analytics
– Social Network analysis
– Text mining / analysis
– Trend analytics
• Social network analysis
– Social network analysis involves studying the
relationships between media users, organizations,
communities etc.
– Emphasizes on analyzing the users in a
network(nodes), and their connections among
each other(edges).
– Depending on the structure of connections, SNA
can help identify
• Opinion leaders
• Influential users
• Influential user communities etc
– SNA can apply static structure mining or dynamic
structure mining.
– Static structure mining works with snapshots of
data of a social network that is stored within a
specific time period.
• Focus on the estructural regularities of the static
network graph.
– Dynamic structure mining uses dynamic data
that constantly keeps changing with time.
• Focuses on unveiling the changes in pattern of data
with the change in time.
– Some of the mining done in SNA are as follows
• Link prediction
• Community detection
• Influence maximization
• Expert findings
• Prediction of trust and distrust among individuals.
• Link prediction Studies a static snapshot of
the social network at a given time t1, and
predicts the future links of the SN for a future
time T2.
– Used for predicting Possible friends suggestions,
as in FB, LinkedIn etc.
– Helps the user to broaden his social links and
connections , wrt his professional/personal friend
circle.
– Leads to an increase in social networking
activities.
• A generic link prediction framework works a s
follows:
– A static social network is input to the prediction
algorithm
– Algorithm applies either a similarity-based
approach or a learning based approach for
prediction of future links in the social network.
– Similarity-based approach calculates the
similarities of non-connected pair of nodes in a SN
and a score is accordingly assigned for each
non-connected pair.
• Based on the descending order of the similarity score, a
list is made to choose the top-N ranked links from the
list for link prediction
– Learning-based approach uses a classifier that
uses some standard machine-learning models to
assign a binary label (+ve or –ve)
• Positive value indicates there is a better chance of
connectivity between non-connected pairs & negative
value indicates vice-versa.
• Community detection Used too find the
correlation between nodes in the network to
assess the strength of connection between
nodes.
– Formation of intra-communities &
inter-communities.
– A user belonging to the same community is
expected to share similar tastes , likes and dislikes.
– Community detection method helps in prediction
of
• What products a user is likely to buy , which movie a
user is likely to watch, services a user may be
interested in etc.
– Different categories of approaches followed for
community detection are as follows
• Traditional clustering methods
• Link-based clustering methods
• Topic-based methods
• Topic-link based methods
• Influence Maximization Influence
propagation is is the task of choosing a set of
proficient users who can prove to be very
efficient for viral marketing.
– This set of efficient users in a SN is called as the
seed set.
– Seed set is very valuable to target for promotion
or publicity, since these users have the highest
reach of spreading information.
– The Seed users can help the other users to decide
in choosing which movie to watch, brand of dress
to buy, which community to join etc.
– Viral marketing works in the following manner
• Initially use an influence maximization technique to
find a set of few influential users in a SN.
• Next influence these users about the goodness and
usefulness of a product , so that it can create a cascade
of influence of buying the same product by the user’s
friends.
• The users friends will turn pass on the influence onto
their friend list, thereby increasing the number of
potential customers for a business.
• This strategy adapted in
– Political campaigns
– Movie recommendations
– Company publicity etc.
• Expert Finding Social networking sites provide a
pool of experts for certain topics and discussions.
– Expert finding is the task of generating and grouping
experts of a SN based on is/her expertise on certain
topics.
– Main task here is to generate/retrieve a ranked list of
top-N experts who are well conversant on a given
topic.
– An expert finding model, takes as input a group of
users U and a task T which consists of set of skills S.
• It then finds individual user/s from U ,that has atleast one
skill S , belonging to T.
– Used in research communities, to help form
collaborations in SN based on expert domain.
• Prediction of trust and distrust among
individuals
– Social networks growing at an exponential rate,
contributing activities and content to the network.
– Issue of trust and distrust among the connected
users.
– Trusted users spread the right information and
positive effects on a SN.
– Distrusted users pose a threat to the SN, and can
cause disturbance in the SN.
– To build a reputed network that canbe trusted by
users, it is essential to trace such individual users
that have a likelihood of conducting harmful
online activities in near future.
– Important to partition users into Trusted users
and Distrusted users.