0% found this document useful (0 votes)
7 views14 pages

OSN Complete Notes

The document provides comprehensive study notes on Online Social Networks (OSNs), covering their definitions, evolution, data collection methods, and ethical considerations. It details the characteristics of modern OSNs, challenges and opportunities they present, and various techniques for data extraction, including API usage and web scraping. Ethical principles and privacy laws governing social media data collection are also discussed, emphasizing the importance of informed consent and data security.

Uploaded by

ak6603585
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views14 pages

OSN Complete Notes

The document provides comprehensive study notes on Online Social Networks (OSNs), covering their definitions, evolution, data collection methods, and ethical considerations. It details the characteristics of modern OSNs, challenges and opportunities they present, and various techniques for data extraction, including API usage and web scraping. Ethical principles and privacy laws governing social media data collection are also discussed, emphasizing the importance of informed consent and data security.

Uploaded by

ak6603585
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Online Social Networks (OSN)

Complete Study Notes — Unit 1, Unit 2 & Unit 3


Covers all syllabus topics with definitions, explanations, examples & comparison tables

UNIT 1: Introduction to Online Social Networks | T: 8 hrs W: 20

1.1 Definition and Evolution of Online Social Networks (OSNs)


What is an Online Social Network (OSN)?
An Online Social Network (OSN) is a web-based platform or application that enables individuals to create
personal profiles, share content, connect with other users, and interact in a virtual community over the internet.
OSNs are built around three core elements: users (nodes), relationships between users (edges), and content
shared between them (posts, images, videos, comments).

Simple Definition: An OSN is a digital space where people create profiles, share information, and form
communities — like Facebook, Instagram, Twitter, or LinkedIn — all accessible over the internet.

📌 Graph View: In graph theory, an OSN is modeled as a Graph G = (V, E) where V = users (vertices/nodes) and
E = relationships (edges/connections).

Types of Relationships in OSNs


• Undirected / Symmetric: Both users agree to connect. Example: Facebook friendship — both people
must accept.
• Directed / Asymmetric: One user can follow another without mutual agreement. Example:
Twitter/Instagram following — you can follow a celebrity without them following back.
• Weighted Relationships: Connections have strength — users who interact more frequently have
stronger links.

Evolution of Online Social Networks


OSNs did not appear suddenly. They evolved through several distinct phases over more than two decades:

Era Year(s) Key Development


Bulletin Board Systems (BBS) and early email lists — text-only
Pre-Web 1970s–1980s
communities.

[Link] — first true OSN; let users create profiles and


First OSN 1997
list friends.

Community Friendster (2002), MySpace (2003) — focused on music, photos,


2000–2004
Era and social circles.

Professional LinkedIn launched — professional networking, job postings, and


2003
OSN career connections.

Mainstream Facebook (2004) became the dominant global OSN. YouTube


2004–2010
Era (2005) for video sharing. Twitter (2006) for micro-blogging.

Instagram (2010) and Snapchat (2011) — mobile-first,


Mobile Era 2010–2014
photo/video focus. WhatsApp for messaging.

Algorithms curate content feeds. TikTok (2016) introduces short-


Algorithm Era 2015–2020
form video. Stories format popularized by Instagram.

Modern Era 2020–Now Reels, live streaming, social commerce, AI-generated content,
Meta, Threads, BeReal. Focus on creator economy.

Characteristics of Modern OSNs


1. User Profiles: Each user has a personal profile with bio, photo, interests, and connections.
2. Content Sharing: Users post text, images, videos, links, stories, and live streams.
3. Social Graph: A map of connections between users — the foundation of every OSN.
4. Feed & Discovery: Algorithmic or chronological feed showing posts from connections and
recommendations.
5. Engagement Features: Likes, comments, shares, reactions, polls, and direct messaging.
6. Privacy Controls: Users can set who sees their content — public, friends, or private.

1.2 Data Collection from Social Networks


Social network data is enormously valuable for research, business intelligence, public health monitoring, and
political analysis. Understanding how this data is collected is fundamental to OSN studies.

Types of Data Available on Social Networks


• Profile Data: Username, age, gender, location, bio, interests, education, job.
• Interaction Data: Likes, comments, shares, retweets, reactions, mentions, tags.
• Content Data: Posts, images, videos, stories, hashtags, links shared.
• Network Data: Friend lists, follower/following relationships, group memberships.
• Behavioral Data: Login times, browsing patterns, time spent on content, click-through rates.
• Metadata: Timestamps, geotags (location), device type, language.

Methods of Data Collection


7. API-Based Collection: Using official APIs provided by platforms (Twitter API, Facebook Graph API).
Structured, authorized, and rate-limited. Most ethical method.
8. Web Scraping: Automated bots crawl public pages and extract data using tools like BeautifulSoup,
Scrapy, Selenium. Useful when APIs are restricted.
9. Surveys & Questionnaires: Researchers ask users directly about their social media behavior and
experiences.
10. Data Donations: Users voluntarily share their own data exports (e.g., downloading Facebook data and
sharing with researchers).
11. Third-Party Data Brokers: Companies that aggregate and sell anonymized social media data for
commercial use.
12. Mobile SDKs: Apps with embedded code that passively collect usage data with user consent.

⚠️Important: Always check the platform's Terms of Service before collecting data. Unauthorized scraping may
violate legal agreements and privacy laws like GDPR.

1.3 Challenges, Opportunities, and Pitfalls in OSNs


Challenges in OSNs
13. Privacy & Data Security: Users often share personal information unaware of how it may be used or
exposed. Data breaches can leak sensitive information to attackers.
14. Misinformation & Fake News: False information spreads faster than true information on social networks
due to emotional engagement and viral sharing.
15. Cyberbullying & Harassment: Anonymity enables hate speech, personal attacks, and coordinated
harassment campaigns.
16. Content Moderation at Scale: Billions of posts daily make it impossible to manually review all content.
Automated systems make errors.
17. Algorithmic Bias & Filter Bubbles: Recommendation algorithms show users only content they already
agree with, reinforcing biases and creating echo chambers.
18. Addiction & Mental Health: Infinite scroll, likes, and notifications are designed to maximize engagement
— leading to addictive usage patterns and anxiety.
19. Data Overload: Collecting and processing massive amounts of unstructured social data is technically
complex and resource-intensive.

Opportunities in OSNs
20. Business Marketing: Targeted advertising based on user demographics, interests, and behavior. Brands
can reach millions for minimal cost.
21. Public Health Monitoring: Tracking disease outbreaks (e.g., COVID-19) and mental health trends
through social media posts.
22. Education & Knowledge Sharing: Platforms like LinkedIn Learning, YouTube, and Twitter communities
enable global knowledge exchange.
23. Political Engagement: Politicians reach citizens directly; civic movements organize and mobilize via
OSNs.
24. Research & Insights: Researchers analyze trends, public opinion, and human behavior at
unprecedented scale.
25. Community Building: People with rare conditions, niche interests, or shared backgrounds find
communities globally.
26. Crisis Response: Emergency services use social media for real-time disaster updates and coordination.

Pitfalls in OSNs
Pitfalls are hidden or overlooked risks that users and organizations often fall into unknowingly:
• Oversharing: Posting too much personal information (location, daily routine) creates security and stalking
risks.
• Digital Addiction: Excessive use interferes with real-life relationships, productivity, and mental wellbeing.
• False Sense of Security: Believing privacy settings fully protect data — many settings are complex and
poorly understood.
• Reputation Damage: Old posts, screenshots, or out-of-context content can permanently harm personal
or professional reputation.
• Exploitation of Vulnerable Users: Children and elderly users are particularly vulnerable to scams,
manipulation, and inappropriate content.
• FOMO (Fear of Missing Out): Constant comparison to curated 'highlight reels' of others leads to anxiety
and low self-esteem.

Challenges Opportunities Pitfalls

Privacy breaches Business marketing Oversharing personal data

Fake news & misinformation Public health monitoring Digital addiction

Cyberbullying Education & knowledge sharing Reputation damage

Filter bubbles Community building False sense of privacy

Content moderation Political engagement Exploitation of vulnerable users

Data overload Crisis response FOMO and mental health issues


1.4 Social Media APIs for Data Extraction
An API (Application Programming Interface) is a set of rules and protocols that allows one software application to
communicate with another. Social media companies provide APIs so that developers and researchers can access
platform data in a structured, controlled, and authorized way.

API in Simple Terms: An API is like a waiter in a restaurant. You (the developer) tell the waiter (API) what you
want. The waiter goes to the kitchen (the platform's database) and brings back what you asked for — without you
entering the kitchen directly.

Why Use APIs for Data Extraction?


• Authorized Access: Data is accessed with permission from the platform — legal and ethical.
• Structured Data: APIs return data in clean JSON or XML format — easy to process.
• Real-Time Data: Many APIs provide live streams of data (e.g., Twitter Streaming API).
• Historical Data: Some APIs allow querying past posts, trends, and interactions.
• Rate Limiting: APIs control how much data can be collected per hour/day — preventing server overload.

Major Social Media APIs


Platform API Name What It Provides

Tweets, user profiles, trends, follower data, hashtag


Twitter / X Twitter API v2
tracking, sentiment streams.

Public posts, page insights, user profiles (with consent),


Facebook Graph API
group data, ads performance.

Business account posts, follower counts, engagement


Instagram Instagram Graph API
metrics, hashtag data.

Video metadata, comments, channel stats, search


YouTube YouTube Data API
results, trending videos.

Professional profiles, job postings, company pages, post


LinkedIn LinkedIn API
engagement.

Posts, comments, subreddit data, upvotes, user karma


Reddit Reddit API (PRAW)
— widely used in research.

Messages from public channels and bots — used for


Telegram Telegram Bot API
monitoring misinformation.

API Access Levels


• Public / Free Tier: Limited data access — suitable for small projects and individual researchers.
• Academic / Research Tier: Greater historical access for verified researchers. Example: Twitter
Academic API.
• Commercial Tier: Full access for paying business customers — real-time data, higher rate limits.
• Enterprise Tier: Full firehose access — all data in real-time. Available to partners and large corporations.

🔑 Key Terms: Rate Limit = maximum API calls per time period. Endpoint = specific URL that gives a specific type
of data. JSON = the data format returned by most APIs.

UNIT 2: Data Collection and Analysis in OSNs | T: 8 hrs W: 20


2.1 Techniques for Collecting Data from Online Social Media
Collecting data from social media platforms requires selecting the right technique based on the research goal, the
platform's access policies, and ethical constraints. Below are the major techniques:

1. API-Based Data Collection


The most common and ethical method. Platforms provide official APIs to access data in a structured manner.
• Tools: Tweepy (Python library for Twitter), PRAW (Reddit), Facebook Graph API Explorer.
• Data format: JSON (JavaScript Object Notation) — easy to parse and analyze.
• Advantages: Official, structured, authorized, real-time or historical data available.
• Disadvantages: Rate limits restrict volume; some data hidden behind premium tiers; APIs can change or
be revoked.
• Example: Collecting 10,000 tweets about a product launch using Tweepy and analyzing customer
sentiment.

2. Web Scraping
Automated programs (bots) visit web pages, read the HTML structure, and extract specific data — simulating a
human browsing the site.
• Tools: BeautifulSoup (HTML parsing), Scrapy (full scraping framework), Selenium (browser automation
for dynamic content).
• Advantages: Works even when no API exists; can collect large volumes of public data.
• Disadvantages: May violate Terms of Service; platforms use CAPTCHAs and bot detection; data is
unstructured HTML.
• Legal Note: Always check if scraping is permitted. The hiQ Labs v. LinkedIn case (US, 2022) ruled
scraping of public data is generally legal.

3. Survey & Questionnaire Methods


Researchers design structured questionnaires to collect self-reported data from social media users about their
behavior, attitudes, and experiences.
• Tools: Google Forms, SurveyMonkey, Qualtrics.
• Advantages: Direct, controlled data collection; captures opinions not visible in posts.
• Disadvantages: Small sample sizes; self-reporting bias; respondents may not answer honestly.

4. Social Media Monitoring Tools


Third-party platforms that aggregate data from multiple social networks simultaneously.
• Examples: Brandwatch, Hootsuite Insights, Sprout Social, Mention, Talkwalker.
• Use case: Brands monitoring their reputation across Facebook, Twitter, Instagram, and Reddit in one
dashboard.

5. Data Donation & Crowdsourcing


Users voluntarily provide their own data (downloaded from platforms) for research purposes.
• Example: Researchers at universities ask volunteers to share their Facebook data download for
behavioral studies.
• Advantage: Highly accurate first-party data with full user consent.

Method Best For Main Tool Key Limitation

Structured, authorized Rate limits, restricted


API Collection Tweepy, PRAW
data data

Web Scraping Public data without API BeautifulSoup, Scrapy ToS issues, bot
detection

User opinions & Small samples, biased


Surveys Google Forms, Qualtrics
attitudes responses

Costly, surface-level
Monitoring Tools Multi-platform tracking Brandwatch, Hootsuite
data

Deep behavioral
Data Donation Platform data exports Small volunteer groups
analysis

2.2 Ethical Considerations in Social Media Data Collection


Collecting data from social media involves real people. Even if data is publicly available, researchers must
carefully consider ethical responsibilities. Unethical data collection can harm individuals, damage trust, and violate
laws.

Core Ethical Principles


27. Informed Consent: Participants should know their data is being collected and agree to it. Problem: Most
social media data is collected without individual consent — raising serious ethical questions.
28. Privacy & Anonymization: Even if data is public, researchers should anonymize it to protect individuals.
Avoid publishing direct quotes linked to real accounts without permission.
29. Data Minimization: Collect only what is necessary for the research purpose. Do not collect extra data
'just in case.'
30. Transparency: Clearly disclose the purpose of data collection, who will access it, and how it will be used.
31. Non-Maleficence (Do No Harm): Ensure research does not expose users to risk — e.g., outing
anonymous users, targeting vulnerable people, or enabling surveillance.
32. Data Security: Collected data must be stored securely, encrypted, and deleted when no longer needed.
33. Compliance with Laws: Respect GDPR (EU), CCPA (California), India's DPDP Act, and platform Terms
of Service.

Key Privacy Laws Governing Social Media Data


• GDPR (General Data Protection Regulation): EU law (2018). Requires user consent, right to be
forgotten, data portability, and strict security. Heavy fines for violations (up to €20 million or 4% of global
revenue).
• CCPA (California Consumer Privacy Act): US law giving California residents rights over their personal
data including the right to know, delete, and opt-out of data selling.
• India's DPDP Act (2023): India's Digital Personal Data Protection Act — requires consent for data
processing, data localization for sensitive data, and rights for data principals.

The Cambridge Analytica Case — Ethics in Practice


The most famous case of unethical social media data collection. In 2014–2018, Cambridge Analytica harvested
Facebook data of 87 million users without explicit consent through a personality quiz app. This data was used to
build detailed psychological profiles and target political advertising during the 2016 US election and Brexit
referendum.

• What happened: A quiz app collected data not just from users who took it, but also from all their
Facebook friends — without those friends' consent.
• Impact: 87 million profiles harvested. Used to influence elections. Mark Zuckerberg testified before US
Congress. Facebook fined $5 billion.
• Lessons learned: Platform APIs must limit third-party data access. User consent must be explicit. Data
collected for one purpose cannot be reused for another (purpose limitation).
⚠️Ethical Rule of Thumb: Just because data is publicly accessible does NOT mean it is ethically acceptable to
collect and use it. Always ask: Would users be comfortable knowing how their data is being used?

2.3 Data Processing and Cleaning for Analysis


Raw social media data is messy, incomplete, and full of noise. Before any analysis can begin, data must go
through a processing pipeline to make it accurate and usable. This is one of the most time-consuming but critical
steps.

The Social Media Data Processing Pipeline


Step Stage What Happens

1 Data Collection Gather raw data via API, scraping, or surveys.

2 Data Ingestion Load raw data into storage (CSV, database, data lake).
Remove duplicates, fix errors, handle missing values, remove
3 Data Cleaning
spam/bots.

Normalize text (lowercase, remove punctuation), tokenize,


4 Data Preprocessing
remove stop words.

Extract useful signals: sentiment scores, hashtags, named


5 Feature Extraction
entities, network metrics.

6 Data Analysis Apply statistical, NLP, or network analysis methods.

Create charts, graphs, word clouds, network diagrams to


7 Visualization
communicate findings.

8 Interpretation Draw conclusions, validate against hypotheses, write reports.

Common Data Quality Issues in Social Media Data


• Duplicate Posts: Same content posted multiple times (retweets, reposts). Must be deduplicated.
• Spam & Bot Accounts: Fake automated accounts generating noise. Must be detected and removed.
• Missing Data: Some fields (location, age) are optional and often absent. Handle with imputation or
exclusion.
• Encoding Issues: Special characters, emojis, multilingual text cause parsing errors.
• Slang & Abbreviations: Social media language differs from formal text — 'LOL', 'brb', '#IYKYK' need
handling.
• Inconsistent Formats: Dates, usernames, and hashtags appear in many formats across platforms.

Text Preprocessing Steps


34. Lowercasing: Convert all text to lowercase: 'Cloud' → 'cloud'
35. Removing Punctuation & Special Characters: Remove @mentions, URLs, #hashtags symbols, emojis
(or encode them).
36. Stop Word Removal: Remove common words with little meaning: 'the', 'is', 'and', 'a'.
37. Tokenization: Split text into individual words (tokens): 'I love cloud' → ['I', 'love', 'cloud'].
38. Stemming / Lemmatization: Reduce words to their root form: 'running' → 'run'; 'better' → 'good'
(lemmatization).
39. Normalization: Standardize slang: 'gonna' → 'going to'; 'gr8' → 'great'.
Analysis Techniques Applied to Social Media Data
• Sentiment Analysis: Classify text as positive, negative, or neutral. Used for brand monitoring, political
analysis, product reviews.
• Network Analysis: Study the social graph — find influencers, communities, information flow paths.
• Topic Modeling: Discover hidden themes in large text corpora using LDA (Latent Dirichlet Allocation).
• Trend Detection: Identify rising topics using hashtag frequency, time-series analysis.
• Bot Detection: Machine learning models trained to distinguish human vs. automated accounts.
• Geospatial Analysis: Map content to geographic locations using geotagged posts.

2.4 Case Studies on Social Media Data Collection

Case Study 1: Cambridge Analytica (2014–2018)


• Context: Political consulting firm harvested Facebook data of 87 million users.
• Method: Personality quiz app exploited Facebook's API to also collect friends' data.
• Use: Built psychological profiles for targeted political advertising (2016 US election, Brexit).
• Outcome: Facebook fined $5 billion. Stricter API access policies implemented. GDPR enacted.
• Lesson: Consent must be explicit. Data purpose must be disclosed. APIs must limit access scope.

Case Study 2: Twitter Data for COVID-19 Research (2020)


• Context: Researchers worldwide used Twitter's COVID-19 research API to study the pandemic.
• Method: Collected millions of tweets with keywords: #COVID19, #lockdown, #vaccine.
• Use: Tracked misinformation spread, public compliance with health guidelines, vaccine hesitancy.
• Outcome: Studies published in Nature, WHO reports. Helped shape public health communication.
• Lesson: Social media data can provide real-time public health intelligence faster than traditional surveys.

Case Study 3: Instagram & Youth Mental Health (2021)


• Context: Facebook's internal research (leaked by whistleblower Frances Haugen) showed Instagram
harms teen girls' mental health.
• Method: Internal A/B testing and user surveys on 14–17 year old female users.
• Use: Showed Instagram's algorithm amplified body image content leading to depression and eating
disorders.
• Outcome: US Congress hearings on social media regulation. Age verification debates. Algorithm
transparency calls.
• Lesson: Platforms have a responsibility to audit the impact of their algorithms — especially on vulnerable
users.

UNIT 3: Trust, Credibility, and Reputation in Social Systems | T: 8 hrs W: 20

3.1 Understanding Trust and Credibility in Online Communities


What is Trust in Online Communities?
Trust in OSNs is a user's belief that another user, platform, or piece of content is reliable, honest, competent, and
benevolent. Trust is the foundation of all social interaction — without it, users would not share information, make
purchases, or rely on online reviews.
Trust: A psychological state where one party is willing to be vulnerable to another party's actions, based on
positive expectations about their behavior.

Types of Trust in OSNs


• Interpersonal Trust: Trust between two users. Example: Trusting a friend's product recommendation on
Facebook.
• Institutional Trust: Trust in the platform itself. Example: Believing Facebook properly protects your data.
• Content Trust: Trusting that a news article, review, or post is accurate and not misleading.
• System Trust: Trusting the algorithms and technical infrastructure to work fairly and reliably.

Factors That Build Trust


40. Consistent Behavior: Users who regularly post accurate, helpful content earn trust over time.
41. Transparency: Platforms and users who are open about their identity, methods, and intentions are more
trusted.
42. Social Proof: High engagement (likes, shares, comments) signals that many others find content
trustworthy.
43. Reputation Scores: Accumulated ratings and reviews create a track record of trustworthiness.
44. Verified Identity: Official verification badges (e.g., Twitter blue tick, LinkedIn verified) signal authenticity.
45. Community Endorsement: Recommendations from trusted community members transfer trust.

What is Credibility?
Credibility refers to the quality of being trusted and believed. In OSNs, credibility applies to both the source (who
is posting) and the content (what is posted).

Dimension Source Credibility Content Credibility

Is the author trustworthy and Is the information accurate and


Definition
knowledgeable? verifiable?

Verified badge, expertise, past Citations, evidence, consistency


Indicators
accuracy with facts

A professor posting about climate A news article citing peer-


Example
science reviewed studies

Misinformation, manipulated
Threats Fake accounts, impersonation
images

Challenges to Trust & Credibility


• Anonymous Accounts: Users can hide behind fake names and avatars, making accountability difficult.
• Deepfakes: AI-generated fake videos or audio of real people erode trust in all media.
• Platform Manipulation: Purchasing fake followers or engagement inflates perceived credibility.
• Algorithmic Amplification: Algorithms promote outrage and sensational content regardless of accuracy.
• Echo Chambers: Users only see content from like-minded sources — reinforcing existing beliefs rather
than testing them.

3.2 Reputation Systems and Their Impact on User Behavior


A Reputation System is a mechanism that collects, aggregates, and displays feedback about users or entities to
help others make trust decisions. Reputation systems are the backbone of trust in online marketplaces, review
platforms, and social networks.
Reputation System: A system that computes a reputation score based on aggregated feedback (ratings,
reviews, endorsements) and uses it to signal trustworthiness to other users.

How Reputation Systems Work


46. Feedback Collection: Users rate each other after interactions (e.g., 5-star rating after eBay purchase).
47. Aggregation: Individual ratings are combined into an overall score (average, weighted, Bayesian).
48. Display: The score is shown prominently to help others decide whether to trust the user/product.
49. Incentive Creation: High reputation = more business/influence. Low reputation = fewer interactions.

Real-World Examples of Reputation Systems


Platform Reputation Mechanism Impact on Behavior

Seller feedback
eBay (positive/neutral/negativ Buyers prefer high-rated sellers. Fraud reduced.
e %)
5-star product + seller
Amazon Purchase decisions heavily influenced by ratings.
ratings with reviews

Host & guest mutual


Airbnb Hosts clean properties; guests behave respectfully.
ratings

Driver & rider ratings


Uber / Ola Poor-rated drivers/riders can be removed.
after each trip

Karma points for helpful High-karma users gain privileges; quality content
Stack Overflow
answers rewarded.

Upvotes/downvotes +
Reddit Popular content rises; low-quality content buried.
karma

Endorsements & Professional credibility established through peer


LinkedIn
recommendations validation.

Followers, retweets,
Twitter / X High-follower accounts treated as more credible.
engagement

Impact of Reputation Systems on User Behavior


• Incentivizes Quality: Users produce better content, products, and services to maintain high ratings.
• Self-Regulation: Community members police each other — reporting low-quality or harmful content.
• Trust Transfer: A stranger's positive reputation makes other users willing to interact with them.
• Gatekeeping: New users without reputation face barriers — chicken-and-egg problem.
• Gaming & Manipulation: Users find ways to artificially inflate scores (fake reviews, purchased followers,
review bombing).
• Fear of Negative Ratings: Sellers, drivers, and freelancers may compromise principles to avoid bad
reviews.

Limitations of Reputation Systems


• Fake Reviews: Businesses pay for positive reviews, skewing scores. Amazon removed 200M+ fake
reviews in 2022.
• Recency Bias: Old poor performance may no longer reflect current quality but still drags down scores.
• Cultural Differences: Rating behavior varies by culture — some cultures rarely give 5-star ratings; others
rarely give low ratings.
• Herding Effect: Users simply copy others' ratings rather than forming independent judgments.
3.3 Detecting Misinformation and Fake News in OSNs
Definitions
Misinformation: False or inaccurate information shared without intent to deceive. People share it believing it is
true.
Disinformation: Deliberately false information created and spread to deceive or manipulate. Intent to deceive is
present.
Malinformation: True information used maliciously — e.g., leaking private conversations to embarrass someone.
Fake News: Fabricated news stories designed to mimic legitimate journalism, often with political or financial
motivation.

Why Does Misinformation Spread So Fast on OSNs?


50. Emotional Triggering: False stories are often sensational, fear-inducing, or outrage-provoking —
emotions drive sharing.
51. Algorithmic Amplification: Engagement-maximizing algorithms promote emotional content — false
news gets more engagement than true news.
52. Speed over Accuracy: Users share quickly without verifying — correction comes too late.
53. Echo Chambers: Content is shared within like-minded groups where it is not questioned.
54. Lack of Media Literacy: Many users cannot distinguish credible sources from fabricated ones.
55. Bot Networks: Automated accounts amplify false content to make it appear popular.

📊 Research Finding: A 2018 MIT study found that false news spreads 6 times faster than true news on Twitter,
reaches more people, and penetrates deeper into social networks.

Methods for Detecting Misinformation & Fake News


56. Automated Fact-Checking (AI/NLP): Machine learning models trained on labeled datasets classify
claims as true, false, or unverified. Tools: ClaimBuster, FullFact AI, Google Fact Check Tools.
57. Source Credibility Analysis: Check the reputation and history of the publishing domain. Analyze author
credentials and past accuracy.
58. Cross-Reference Verification: Check whether the claim is reported by multiple independent, reputable
news sources.
59. Network Propagation Analysis: Examine HOW the content spreads — bot-driven or organic?
Coordinated inauthentic behavior detection.
60. Image & Video Verification: Reverse image search (Google Images, TinEye) to check if an image is real
or taken out of context. Tools for deepfake detection.
61. Temporal Analysis: Check when the content was originally created vs. when it is being shared — old
images/videos resurface in new contexts.
62. Crowdsourced Fact-Checking: Platforms like Twitter's Community Notes (formerly Birdwatch) let
trusted users add context to misleading tweets.

Major Fake News Detection Tools & Platforms


Tool / Platform Type How It Works

[Link] Manual fact-checking Journalists investigate and rate viral claims.

Focuses on US political claims; rates


[Link] Manual fact-checking
accuracy.

Google Fact Check Tools AI + manual Aggregates fact-checks; labels results in


search.

Automatically identifies check-worthy claims in


ClaimBuster AI/NLP
text.

Trusted users add context notes to misleading


Twitter Community Notes Crowdsourced
tweets.

Shows how misinformation spreads through


Hoaxy Network visualization
Twitter.

InVID / WeVerify Media verification Verifies authenticity of videos and images.

3.4 Methods for Enhancing Trustworthiness in Social Media


Increasing trustworthiness on social media is a shared responsibility — platforms, governments, researchers, and
users all play a role.

Platform-Level Methods
63. Content Moderation: AI + human review teams flag and remove harmful, false, or violating content.
Example: Facebook's AI removes millions of fake accounts daily.
64. Verified Badges: Blue checkmarks or verification badges signal that an account is genuine. Example:
Instagram, LinkedIn, YouTube.
65. Transparency Reports: Platforms publish regular reports on content removal, data requests, and
moderation decisions.
66. Algorithm Transparency: Explaining how content is recommended — reducing perception of hidden
manipulation.
67. Reducing Virality of Unverified Content: Twitter adds friction: 'Would you like to read this article before
sharing?' WhatsApp limits forwarding to 5 chats.
68. Age Verification: Protecting minors from harmful content and manipulation.

AI & Technology-Based Methods


69. Automated Fake Account Detection: ML models detect bot accounts based on posting patterns,
follower ratios, and profile completeness.
70. Deepfake Detection: AI tools (Microsoft Video Authenticator, Deepware Scanner) identify AI-generated
fake media.
71. Hate Speech Detection: NLP models flag and filter hate speech, harassment, and threatening language.
72. Spam Filtering: ML-based systems block malicious links, phishing content, and spam messages.

Policy & Legal Methods


• GDPR / Data Protection Laws: Force platforms to be transparent about data use and give users control.
• EU Digital Services Act (DSA) 2022: Requires large platforms to conduct risk assessments, be
transparent about algorithms, and cooperate with researchers.
• Platform Terms of Service Enforcement: Consistent enforcement of rules against fake accounts, spam,
and harmful content.
• Government Oversight & Regulation: Parliamentary inquiries, regulatory fines for inaction on harmful
content.

User-Level Methods
73. Media Literacy Education: Teaching users to critically evaluate sources, check facts, and identify
manipulative content.
74. Privacy Settings Management: Users should review and configure who sees their content and personal
data.
75. Two-Factor Authentication: Secures accounts against unauthorized access — protects identity.
76. Reporting Mechanisms: Using platform tools to report fake accounts, misinformation, and harmful
content.
77. Following Credible Sources: Curating a feed of verified, reputable accounts reduces exposure to
misinformation.

Level Method Example

Facebook removes millions of


Platform Content moderation
fake accounts daily

Platform Verified badges Twitter/X blue verification mark

ML models flag coordinated


AI/Tech Bot detection
inauthentic behavior

AI/Tech Deepfake detection Microsoft Video Authenticator

Right to be forgotten, consent


Policy/Law GDPR compliance
requirements

Algorithm transparency for large


Policy/Law EU Digital Services Act
platforms

Fact-checking before sharing viral


User Media literacy
content

User Two-factor authentication Protect accounts from hijacking

Quick Revision Summary — All 3 Units


Topic Key Points to Remember

Web-based platforms for profile creation, content sharing, and social


OSN Definition
connection. Modeled as graphs.

SixDegrees (1997) → Friendster, MySpace → Facebook, YouTube, Twitter


OSN Evolution
→ Instagram, Snapchat → TikTok → Modern AI-driven platforms.
APIs (most ethical), Web Scraping, Surveys, Monitoring Tools, Data
Data Collection Methods
Donation.

Privacy, fake news, cyberbullying, filter bubbles, addiction, content


Challenges in OSNs
moderation at scale.

Marketing, healthcare monitoring, education, community building, political


Opportunities in OSNs
engagement, research.

Oversharing, digital addiction, FOMO, reputation damage, exploitation of


Pitfalls in OSNs
vulnerable users.

Provide authorized, structured access to platform data. Examples: Twitter


Social Media APIs
API, Facebook Graph API, YouTube Data API.

Rate limit (calls per hour), Endpoint (specific data URL), JSON (data
API Key Terms
format returned).

Informed consent, privacy, data minimization, transparency, do no harm,


Ethical Principles
legal compliance.

GDPR EU law for data privacy. Consent required, right to be forgotten, heavy
fines for violations.

87M Facebook profiles harvested without consent. Used for political


Cambridge Analytica
profiling. Led to stricter API policies.

Remove duplicates, handle missing values, remove spam, normalize text,


Data Cleaning Steps
tokenize, remove stop words.

Lowercase → remove punctuation → stop words → tokenize →


Text Preprocessing
stem/lemmatize.

Belief in reliability of users, platforms, and content. Built by consistency,


Trust in OSNs
transparency, and reputation.

Source credibility (who posts) + Content credibility (what is posted). Both


Credibility
must be verified.

Collect feedback → aggregate into score → display to users. Examples:


Reputation Systems
eBay, Amazon, Airbnb, Reddit.

Misinformation = false but not intentional. Fake News = deliberately


Fake News vs Misinformation
fabricated. Spreads 6x faster than truth (MIT).

AI/NLP classifiers, source analysis, cross-reference, reverse image search,


Fake News Detection
network analysis, Community Notes.

Content moderation, verified badges, algorithm transparency, GDPR, AI


Enhancing Trust
bot detection, media literacy education.

★ End of Online Social Networks Notes — Unit 1, 2 & 3 ★


Prepared for Mid-Semester Examination

You might also like