0% found this document useful (0 votes)
7 views37 pages

Essential Data Collection Methods

The document provides an overview of data collection in the data science cycle, emphasizing the importance of systematic methodologies for gathering accurate and reliable data. It discusses various types of data, methods of collection such as experiments, observations, surveys, and transactional data, and highlights the significance of sampling techniques and potential errors like sampling bias and measurement error. Additionally, it covers modern approaches like web scraping and social media data collection, along with practical applications and examples.

Uploaded by

Emmanuel
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views37 pages

Essential Data Collection Methods

The document provides an overview of data collection in the data science cycle, emphasizing the importance of systematic methodologies for gathering accurate and reliable data. It discusses various types of data, methods of collection such as experiments, observations, surveys, and transactional data, and highlights the significance of sampling techniques and potential errors like sampling bias and measurement error. Additionally, it covers modern approaches like web scraping and social media data collection, along with practical applications and examples.

Uploaded by

Emmanuel
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Data Collection

Introduction
• Data collection is the first step in the data science
cycle. It involves systematically gathering the
necessary data to meet a project's objectives.
• With today’s ever-increasing volume of data, a
robust approach to data collection is crucial for
ensuring accurate and meaningful results. This
process requires following a comprehensive and
systematic methodology designed to ensure the
quality, reliability, and validity of data gathered
for analysis.
Types of Data
• Primary Data: Collected directly from the source
through experiments, surveys, interviews, or
observations.
• Secondary Data: Obtained from existing sources
like databases, research papers, government reports,
or online repositories.
• Let's watch a video Tutorial
Data Collection Methods
Overview of Primary Data
Collection Methods
• Data collection refers to the systematic and well-
organized process of gathering and accurately
conveying important information and aspects related
to a specific phenomenon or event.
• Additionally, it is important to take note of the
environment and geographic location from where
the data was obtained, as it can significantly influence
the decision-making process and overall conclusions
drawn from the data.
• Before collecting data, it is essential for a data
scientist to have a clear understanding of the
project’s objectives, which involves identifying the
research question or problem
Common Data Collection Methods

• Data collection can be carried out through various


methods, depending on the nature of the research or
project and the type of data being collected.
• Some common methods for data collection include
i. Experiments
ii. Observation
iii. Transaction
iv. Surveys,
v. Focus groups (Lets watch a video on this approach),
vi. Interviews- Interviews vs Focus groups,
vii. Document analysis.
#1. Collecting data Trough
Experiments
• Data collection through experiments means actively
generating data by manipulating variables in a
controlled setting and observing the outcomes.
• It involves manipulating one or more independent
variables (IVs) to observe their effect on one or
more dependent variables (DVs), while controlling
other factors to ensure that observed changes are
due to the manipulation and not external influences.-
Lets watch the below tutorials
i. Video 1 – What are the characteristics of
experiments?
ii. Video 2
Principles/Characteristics of
Experiments
• Lets watch the below tutorial
i. Video 1
Types of Experimental Designs

i. True Experimental Design


ii. Quasi-Experimental Design
iii. Pre-Experimental Design
• True experiments include random assignment,
control groups, and manipulation of independent
variables. These are the most reliable designs for
establishing cause-and-effect relationships e.g. Rising
global temperatures, Glaciers melt more rapidly.
• Quasi-Experimental-These designs lack random
assignment but still include manipulation of variables.
They are practical for real-world settings like schools,
hospitals, or communities where randomization isn’t
possible.
Types of Experimental Designs

• Pre-Experimental-These are the most basic forms


of experimental design. They often lack
randomization and control groups, making them less
reliable for establishing causality. However, they are
useful for preliminary studies or when full control
isn’t possible.
• Lets watch this video
#2 Observational Data Collection

• Observation is a data collection technique in which


the researcher systematically watches and records
what people/subjects do, how they behave, or how
events unfold in a particular context.
• Lets watch this video
#3 Transactional Data

• Transactional data is generated from day-to-day


operations of an organization and recorded in
systems like point-of-sale (POS) systems, ERP, AMS,
CRM, or e-commerce platforms.
• Transactional data is collected by directly recording
transactions that occur in a particular setting,
such as a retail store or an online platform that
allows for accurate and detailed information on
actual consumer behavior.
• It can include financial data, but it also includes data
related to customer purchases, website clicks,user
interactions, or any other type of activity that is
recorded and tracked.
#4 Surveys

• Survey data is information collected directly from


people through questionnaires or interviews,
where respondents provide answers about their
opinions, behaviors, or characteristics.
• Surveys help data scientists gather primary data ,
information that didn’t exist before
Steps in Survey-Based Data
Collection
Designing Effective Survey
Questions
• A well-designed survey produces clean, analyzable
data. Poorly designed questions lead to confusion
and bias.
Sampling

• Because surveying an entire population is often


impossible, you’ll collect data from a sample a
representative subset.
• Sampling methods are broadly divided into two
categories:
– Probability Sampling- Each individual in the
population has a known and non-zero chance of
being selected.
– Non-Probability Sampling-Selection is not
random ,some members have a higher chance
of being chosen due to convenience or
researcher judgment.
Probability Based Sampling
Techniques

• Lets watch a video tutorial


Non-Probability Based Sampling
Techniques

• Lets watch a video tutorial


Sampling Error
• Sampling error is the difference between the
results obtained from a sample and the true
value of the population parameter it is intended
to represent.
• It is caused by chance and is inherent in any
sampling method. The goal of researchers is to
minimize sampling errors and increase the accuracy
of the results.
• To avoid sampling error, researchers can increase
sample size, use probability sampling methods,
control for extraneous variables, use multiple
modes of data collection, and pay careful
attention to question formulation.
Sampling Error Example
• The supermarket surveys 400 customers about
satisfaction with checkout [Link] true
population average satisfaction is 4.2/5, but the
sample average comes out to 4.0/[Link] difference
(0.2) is the sampling error.
Sampling Bias
• Sampling bias occurs when the sample used in a
study isn’t representative of the population it
intends to generalize to, leading to skewed or
inaccurate conclusions.
• This bias can take many forms, such as selection
bias, where certain groups are systematically over-
or underrepresented, or volunteer bias, where only
a specific subset of the population participates.
• A supermarket only surveys customers during
weekday mornings, missing working professionals
who shop in the [Link] results suggest high
satisfaction, but they exclude a key segment —
causing sampling bias.
Measurement Error
• A measurement error happens when the data
collected is inaccurate the responses or recorded
values don’t reflect the true situation.
• This can occur during data collection, recording, or
interpretation.
• Example: How satisfied are you with the speed of
the checkout and staff friendliness?” Customers are
confused because it combines two ideas checkout
speed and staff behavior so their answers are
inconsistent.
Summary of errors
A Sampling Case Study
Survey Data Collection Tools
Web Scraping Data Collection

• Web scraping and social media data collection are


two approaches used to gather data from the
internet. Web scraping involves pulling information
and data from websites using a web data extraction
tool, often known as a web scraper.
• An example would be a travel company looking to
gather information about hotel prices and
availability from different booking websites. Web
scraping can be used to automatically gather this data
from the various websites and create a
comprehensive list for the company to use in its
business strategy without the need for manual work.
Using Python to Scrape Data from
the Web
• Python is one of the popular programming
languages used for web scraping due to its various
libraries and frameworks that make it easy to pull and
process data from websites.
• Python Libraries for Web scraping
Using Python to Scrape Data from
the Web
• The first step of web scraping is to find a table we
want to scrape, which means figuring out the table
and web page we want to scrape. E.g scraping PAYE
rates in Kenya from this website
– Check the table we want to scrape. First right-click
on the table. Now it will show us an ‘inspect’
option
– The next step is to click the inspect option. It will
open the HTML document of that specific web
page.
– Next, we identify the class of our table before
we can start scraping
Using Python to Scrape Data from
the Web
• Let's review the web scraping example on google
colab
Social Media Data Collection

• Social media data collection involves gathering


information from various platforms like Twitter and
Instagram using application programming interface or
monitoring tools.
• An application programming interface (API) is a set
of protocols, tools, and definitions for building software
applications allowing different software systems to
communicate and interact with each other and enabling
developers to access data and services from other
applications, operating systems, or platforms
• Challenges of Social Media API Data Collection
include- Access Restrictions and Authentication, Rate
Limits and Data Volume, Data Incompleteness,
Changing API Policies, Ethical and Legal Considerations
Common Social Media API’s

• APIs provided by social media platforms e.g


LinkedIn REST API allow data scientists to collect
structured data on user interactions and content.
• Facebook / Meta APIs include Graph API,
Instagram Graph API, WhatsApp Business API
• X (formerly Twitter) APIs - Twitter API v2
• YouTube APIs - YouTube Data API, YouTube
Analytics API
• Tiktok APIs -TikTok for Developers API, TikTok
Marketing API
Snscrape library

• snscrape (short for social-network-scraper) is a


Python library used to extract public posts,
profiles, and metadata from popular social media
platforms without needing an API key.
• It works by scraping publicly available data directly
from the web interface.
Social Listening and Network
Analysis
• Social listening involves monitoring online
conversations for insights on customer behavior and
trends. Surveys conducted on social media can
provide information on customer preferences and
opinions.
• Network analysis, or the examination of
relationships and connections between users, data, or
entities within a network, can reveal influential users
and communities. It involves identifying and
analyzing influential individuals or groups as well as
understanding patterns and trends within the
network.
Examples of Social Media Data uses
• An example of social media data collection is conducting a Twitter
survey on customer satisfaction for a food delivery company.
Data scientists can use Twitter's API to collect tweets containing
specific hashtags related to the company and analyze them to
understand customers' opinions and preferences.
– Analyzing Public Sentiment on Political Elections Using Twitter
Data
– Machine Learning Approaches for Emotion Detection in
Instagram Comments
– Detecting Fake News on Twitter Using Natural Language
Processing Techniques
– An AI-Based Framework for Identifying Misinformation During
Pandemics
– Mining Twitter Data for Early Detection of Mental Health Issues
– AI-Based Approaches for Identifying Cyberbullying in Online
Communities
Practical

• Work on the Practical Exercise on web scraping and


social media data collection provided on the e-learning
Class Activity

1. A café wants to understand why repeat visits have


declined over the past three months. Which data
collection method would be most appropriate and why?
2. A farming company wants to determine if a new
fertilizer increases crop yield compared to the current
fertilizer. Which method should they use?
3. A company suspects that employees are experiencing
high stress due to workload and tight deadlines. What
data collection method would provide honest and
detailed insights?
4. The police want to know how safe residents feel in their
neighborhood and what issues worry them most. What
method is suitable?
5. Environmental scientists need accurate and continuous
data about pollution levels in a river. What method
should they use?
Thank you!

Any Questions?

You might also like