0% found this document useful (0 votes)
3 views16 pages

Chapter 2

Chapter 2 of the Data Analysis course focuses on data collection, defining its importance and various methods including primary and secondary data collection techniques. It covers methods such as questionnaires, experiments, and observations, as well as sampling techniques and their significance in research. The chapter emphasizes the need for careful data collection to ensure accurate research outcomes.

Uploaded by

lilimeriem20
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views16 pages

Chapter 2

Chapter 2 of the Data Analysis course focuses on data collection, defining its importance and various methods including primary and secondary data collection techniques. It covers methods such as questionnaires, experiments, and observations, as well as sampling techniques and their significance in research. The chapter emphasizes the need for careful data collection to ensure accurate research outcomes.

Uploaded by

lilimeriem20
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Data Analysis course

Chapter 2: Data Collection

Contents
1 Introduction 1

2 What is Data Collection? 1

3 Data Collection Methods 2


3.1 Primary Data Collection . . . . . . . . . . . . . . . . . . . . . . . 2
3.1.1 Questionnaires/Surveys . . . . . . . . . . . . . . . . . . . 2
3.1.2 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . 4
3.1.3 Observations . . . . . . . . . . . . . . . . . . . . . . . . . 5
3.2 Secondary Data Collection: . . . . . . . . . . . . . . . . . . . . . 6
3.2.1 Databases . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
3.3 Sampling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8
3.3.1 Sampling Techniques . . . . . . . . . . . . . . . . . . . . . 8
3.3.2 Sampling Bias . . . . . . . . . . . . . . . . . . . . . . . . 13

1 Introduction
Analysis relies on data, which provides the information needed to identify pat-
terns, explain phenomena, and make informed decisions. This chapter defines
the data collection process, highlighting various data collection methods and
sampling techniques.

2 What is Data Collection?


Data collection is the process of gathering information from different sources to
address a specific problem or research question. Data collection as a main stage
can overshadow the quality of achieving results by decreasing the possible errors
which may occur during a research project. Therefore, alongside a good design
for the study, plenty of quality time should be spent in the collection of data to
gain appropriate results since insufficient and inaccurate data prevents assuring

1
the accuracy of findings [1]. On the other hand, although a suitable data col-
lection method helps to plan good research, it cannot necessarily guarantee the
overall success of the research project[3].

3 Data Collection Methods


Generally, data collection methods are divided to two main categories of Primary
Data Collection Methods and Secondary Data Collection Methods [4] [2].

3.1 Primary Data Collection


Primary data is fresh and original, collected directly from first-hand sources and
never used before. The information obtained through primary data collection
techniques is precise and tailored to the purpose of the study.

3.1.1 Questionnaires/Surveys
• A questionnaire/survey is a research instrument consisting of a series
of questions for the purpose of gathering information from respondents.
Questionnaires can be thought of as a kind of written interview.
• They can be carried out face to face, by telephone, post or online.
• Questionnaires/surveys can be an effective means of measuring the behav-
ior, attitudes, preferences, opinions and, intentions of a large numbers of
respondents more cheaply and quickly than other methods.
• A questionnaire/survey can be structured (i.e., using only closed-ended
questions), unstructured (i.e., using only open-ended questions) or semi-
structured (i.e., using both closed-ended and open- ended questions).
Open-ended questions Closed-ended questions

– Open questions allow people to – Closed questions allow


express what they think in as multiple-choice answers that the
much detail as they like in their researcher specifies.
own words. – These multiple-choice answers
– E.g. “Can you tell me what you can be dichotomous (i.e., two
think about your English choices) or polytomous (i.e., more
language syllabus?” than two choices).
– ........................................ – These multiple-choice questions
– ........................................ can consist of a list of responses
or a continuous rating scale.

• Example: Online Classes / Education

2
– Open-ended: “What challenges have you faced with online learn-
ing?”
– Closed-ended: “On a scale of 1–5, how effective do you find online
classes compared to in-person classes?”
• Survey/Questionnaire Design

1. Define the Purpose & Objectives


– What do I want to learn from this survey?*
– Example: “I want to measure student satisfaction with online
classes.”
– Keep objectives specific and measurable. E.g. “I want to know
everything about students.” (too vague).
2. Identify the Target Population
– Who will answer the survey? (students, patients, employees,
customers, etc.)
3. Choose the Type of Survey (Structured, Unstructured, Semi-structured)
4. Draft the Questions
– Keep Your Questionnaire Short: Respondents are less likely to
answer a long questionnaire than a short one, and often pay
less attention to questionnaires which seem long, monotonous,
or boring
– Decide the format (Open-ended, Closed-ended, Rating scales,
Multiple choice)
– Avoid technical terms and jargon. Words used in surveys should
be easily understood by anyone taking the survey.
– Avoid Vague or Imprecise Terms. Usually, it’s best to use terms
that will have the same specific meaning to all respondents.
– Define Things Very Specifically: For example, don’t ask: “What
is your income?” A better question would be specific and might
ask: “What was your total household income before taxes in
2005?”
– Avoid Complex Sentences: Sentences with too many clauses or
unusual constructions often confuse respondents. Scales that ask
respondents to make complex calculations can cause problems.
– Provide Reference Frames: Make sure all respondents are an-
swering questions about the same time and place. For example,
if you ask: “How often do you feel sad?” some people might pro-
vide an answer about their life’s experience, while others might
only be thinking about today. Usually, it’s better to provide a
reference frame: “How often have you felt sad during the past
week?”

3
– Make Sure Scales Are Ordinal: If you are using a rating scale,
each point should be clearly higher or lower than the other for
all people. For example, don’t ask “How many jobs are available
in your town: Many, a lot, some, or a few. “ It’s not clear to
everyone that “a lot” is less than “many.” A better scale might
be: “A lot, some, only a few, or none at all.”
– Avoid Double-Barreled Questions. Questions should measure
one thing. Double barreled questions try to measure two (or
more!) things.
– Answer Choices Should Anticipate All Possibilities. If a respon-
dent could have more than one response to a question, it’s best
to allow for multiple choices. If the categories you provide don’t
anticipate all possible choices, it’s often a good idea to include an
“Other-Specify” category. If You Want a Single Answer, Make
Sure Your Answer Choices Are Unique and Include all Possible
Responses If you are measuring something that falls on a contin-
uum, word your categories as a range.
5. Organize the Questionnaire
– Start a questionnaire with an introduction and provide a title for
each section.
– Start with easy, general questions (to engage the respondent).
– Move to specific or sensitive questions later.
– Group similar topics together (e.g., satisfaction, behavior, demo-
graphics).
6. Pilot Test the Survey
– Try it with a small sample first.
– Check: Are questions clear? Are responses useful? Is it too long?
– Revise accordingly.
7. Finalize and Distribute
– Choose the medium: online (Google Forms, SurveyMonkey),
paper-based, phone, or in-person.
– Ensure anonymity/confidentiality if needed.

• Survey/Questionnaire Implementation
See this link

3.1.2 Experiments
• Experimental studies involve manipulating variables to observe their im-
pact on the outcome. Researchers control the conditions and collect data
to conclude cause-and-effect relationships.
• Key Principles of Experimental Data Collection

4
– Independent Variable: The factor that the researcher intention-
ally changes or manipulates.
– Dependent Variable: The variable that is measured to see if it
changes as a result of the independent variable.
– Controlled Environment: Experiments are conducted in a setting
where other potential influencing factors are minimized to ensure the
observed effect is due to the independent variable.
• Example: Marketing & Advertising
– Goal: Testing the effectiveness of an ad campaign.
– Independent Variable: Type of advertisement (video ad vs. banner
ad vs. social media post).
– Dependent Variable: Number of clicks, purchases, or brand recall.
– Controlled Environment: Same product, same target audience, same
time frame; only the ad format changes.

• Example: Studying the effect of a new treatment on patient recovery.


– Independent Variable: Type of treatment (new drug vs. standard
drug).
– Dependent Variable: Patient recovery time or improvement in symp-
toms.
– Controlled Environment Controlled Environment: Patients of similar
age and health condition, same hospital environment, monitored over
the same period.

3.1.3 Observations
• Observation is a data collection method that involves the direct observa-
tion of events, people, units or phenomena in their natural setting. It aims
to understand how participants interact/behave. Observation is typically
divided into four types:

– Non-participant (or naturalistic) observation is used extensively


in case study research in which the researcher tries to observe with
no intervention events, activities, and interactions in order to gain a
direct understanding of a phenomenon in its natural context.
Examples:
∗ Observing how patients behave in a waiting room (levels of stress,
interactions with staff) without intervening.
∗ Watching how shoppers move through a supermarket without
interacting with them, to study buying patterns.

5
– Participant observation involves the researcher’s intervention in the
environment. In other words, the researcher joins a group as a par-
ticipating member to get a first-hand perspective of the group and
their activities.
Examples:
∗ A researcher works alongside factory workers to understand the
challenges they face with machines or processes.
∗ A researcher teaches part of a class while observing how students
interact with the lesson.
– Structured observation involves going into the classroom with
a specific focus, with concrete observation categories. It involves
completing ‘an observation scheme’ or ‘observation checklist’.
Examples:
∗ Using a checklist to record how many customers stop at a product
display, pick up the product, or purchase it.
∗ Using an observation rubric to measure student participation
(e.g., number of questions asked, level of group interaction).
– Unstructured observation is less clear on what is looking for and
the researcher needs to observe first what is taking place before de-
ciding on its significance for the research. It involves completing
narrative field notes (exploratory in nature).
Examples:
∗ Writing descriptive notes on classroom dynamics without prede-
fined categories—e.g., noticing how students form study groups

3.2 Secondary Data Collection:


Data that has already been used is known as secondary data. The researcher
has access to data from organizational and external sources.

3.2.1 Databases
Databases store and organize data that has already been collected by others.
This secondary data can include:
• Public Records: Data from government agencies, public databases, and
census records.

• Published Sources: Information from academic journals, industry re-


ports, and other published materials.
• Existing Datasets: Large datasets compiled by research institutions and
other organizations.
• Example

6
– Kaggle datasets on customer behavior, e-commerce, or ad click-through
rates.
– Google Trends data on consumer search interests.
– MIMIC-III dataset (critical care patient data).
– UCI Machine Learning Repository (pneumonia, breast cancer, dia-
betes datasets).
• Python Implementation: To analyze such databases, we first need to im-
port them into Python. Below are common ways:
• Importing CSV / Excel files
import pandas as pd # Import pandas library for data
analysis

# Import CSV file into a DataFrame


df_csv = pd . read_csv ( " data . csv " )

# Import Excel file into a DataFrame


df_excel = pd . read_excel ( " data . xlsx " )

# Display the first 5 rows of the CSV data


print ( df_csv . head () )

• Importing from MySQL


import mysql . connector # Library to connect to MySQL
databases
import pandas as pd # Library for data analysis

# Connect to the MySQL database ( change credentials


accordingly )
conn = mysql . connector . connect (
host = " localhost " ,
user = " your_username " ,
password = " your_password " ,
database = " your_database "
)

# Write and execute SQL query , load data into a


DataFrame
query = " SELECT * FROM my_table "
df = pd . read_sql ( query , conn )

# Show the first 5 rows


print ( df . head () )

# Close the connection


conn . close ()

7
3.3 Sampling
• Sampling is an indispensable technique in research that allows researchers
to infer information about a population based on results from a subset of
the population, without having to investigate every individual member.

• The sample is the selected group of elements/individuals who will actually


participate in the research and be representative of the whole population.
What is the difference between sample and population?
– The population is the entire group that you want to draw conclu-
sions about.
– The sample is the specific group of individuals that you will collect
data from.
• E.g. You are doing research on working conditions at Company X. Your
population is all 1000 employees of the company. Your sampling frame is
the company’s 100 selected employee.

3.3.1 Sampling Techniques


Broadly, there are two approaches to sample research design: (a) Probability
Sampling and (b) Non-probability Sampling.

8
1. Probability Sampling
• In general, with probability sampling, all elements (eg., persons,
houses. . . ) in the population have the chance of being included in
the sample. This technique gives the most reliable representation of
the whole population.
• Probability sample designs includes the following types/strategies:

9
a. Simple Random Sampling: every item in the population has
an equal chance of being included in the sample.
Python Implementation:
# importing the random module
import random

# defining the population from where sample will


be created
population = list ( range (1 , 100) )

# defining the size of sample


sample_size = 10

# perform simple random sampling by using the


random . sample () function
sample = random . sample ( population , sample_size )

# it will print 10 random numbers within the range


provided
print ( " Simple random sampling of 10 numbers are :
" , sample )

# # output :
# # Simple random sampling of 10 numbers are : [40 ,
34 , 60 , 96 , 8 , 95 , 94 , 73 , 93 , 26]

b. Systematic Sampling: every nth case after a random start is


selected.
Python Implementation:
import numpy as np

# Define the population


population = np . arange (1 , 100)

# Define the sample size and sampling interval


# We can provide sample size and sampling
interval as per user
# Here sampling interval is 9 (99//10 = 9)
sample_size = 10
sampling_i nterval = len ( population ) //
sample_size

# Define the starting point of the sample


# refer to https :// numpy . org / doc / stable / reference
/ random / generated / numpy . random . randint . html
for random . randint ()
start_point = np . random . randint (0 ,
sampling_interval )

10
# Perform systematic sampling
sample = population [ start_point ::
sampling_interval ]

# Print the sample


print ( " Systematic or interval sampling of 10
numbers are : " , sample )

# # Output :
# # Systematic or interval sampling of 10 numbers
are : [ 6 15 24 33 42 51 60 69 78 87 96]
# # Random numbers are generated from integer 6
with 9 th integer as 15

c. Stratified Sampling: the population is divided into strata (or


subgroups) and a random sample is taken from each subgroup.
Python Implementation:
import pandas as pd
from sklearn . model_selection import
train_test_split

# Load the data into a Pandas DataFrame


data = pd . read_csv ( " https :// raw . githubusercontent
. com / ayan - zz / Statistics_python / main / titanic .
csv " )

# Specify the stratification variable


stratify_by = ’ Sex ’

# Split the data into training and testing sets ,


with stratification
train , test = train_test_split ( data , test_size
=0.3 , stratify = data [ stratify_by ])

# Check the distribution of the stratification


variable in the training and testing sets
print ( " Train dataset :\ n " , train [ stratify_by ].
value_counts () )
print ( " Test dataset :\ n " , test [ stratify_by ].
value_counts () )
# # Output :
# # Train dataset :
# # male 403
# # female 220
# # Name : Sex , dtype : int64

# # Test dataset :
# # male 174

11
# # female 94
# # Name : Sex , dtype : int64

d. Cluster Sampling: the population is divided into subgroups,


but each subgroup should have similar characteristics to the
whole population. Instead of sampling individuals from each
subgroup, you randomly select entire subgroups.
Python Implementation:
import random

# create a list of population data


population_data = [10 , 15 , 20 , 26 , 29 , 33 , 35 ,
37 , 41 , 42 ,
46 , 48 , 50 , 52 , 55 , 58 , 61 ,
64 , 68 , 72]

# set the desired cluster size


cluster_size = 5

# randomly select a starting point for the first


cluster
starting_point = random . randint (0 , cluster_size +
1)

# create a list to store the sampled data


sampled_data = []

# loop through the population data by clusters of


size cluster_size
for i in range ( starting_point , len (
population_data ) , cluster_size ) :
# append the data from the current cluster to
the sampled data list
sampled_data . append ( population_data [ i +1: i +
cluster_size ])

# print the sampled data


print ( starting_point )
print ( sampled_data )
# # Output
# #0
# #[[23 , 26 , 29] , [37 , 41 , 42] , [50 , 52 , 55] , [64 ,
68 , 72]]

2. Non-probability Sampling
• Non-probability sampling, in contrast, refers to the sampling process
in which the samples are selected for a specific purpose with a prede-

12
termined basis of selection. In other words, units of the sample are
chosen on the basis of personal judgment or convenience.
• Non-probability techniques, relying on the judgment of the researcher,
cannot be commonly used to make generalizations about the whole
population.
• Non-probability sample designs include the following strategies:

a. Convenience Sampling: simply includes the individuals who hap-


pen to be most readily and easily accessible to the researcher.
b. Snowball sampling: uses a few cases to help encourage other cases
to take part in the study, thereby increasing sample size. This ap-
proach is most applicable in small populations that are difficult to
access due to their closed nature, e.g. secret societies and inaccessible
professions.
c. Purposive/judgement sampling: selecting a sample that is most
useful to the purposes of the research, it must have clear criteria and
rationale for inclusion.
d. Quota sampling: participants are chosen on the basis of prede-
termined characteristics so that the total sample will have the same
distribution of characteristics as the wider population.

3.3.2 Sampling Bias


Sampling bias refers to the selection of a sample that is not representative of the
population being studied. This can lead to conclusions that are not accurate for
the population as a whole. For example, if a survey is conducted to determine

13
the popularity of a political candidate, but only a small group of people are
selected to participate in the survey, the results may not accurately reflect the
views of the larger population. To reduce sampling bias, it is important to select
a sample that is representative of the population and to use random sampling
methods.

14
• Selection Bias: This occurs when certain data points are systematically
included or excluded from the dataset, resulting in a non-representative
sample. For instance, a healthcare study that only involves healthy pa-
tients may not be applicable to the broader population.
Example: If a survey on adolescent alcohol consumption excludes schools
in disadvantaged neighborhoods, the resulting sample wouldn’t represent
the entire adolescent population. Consequently, the prevalence of alcohol
use in these neighborhoods might be underestimated.
• Survivorship Bias: This bias arises when only surviving or successful
data points are considered, creating an overly positive representation.
Example: Analyzing only adolescents who have survived heavy alcohol
consumption episodes excludes those who haven’t faced severe alcohol-
related issues, possibly underestimating the risks of alcohol use among all
adolescents.
• Non-Response Bias: This occurs when certain segments of the popula-
tion do not respond to surveys or data collection efforts, leading to skewed
results if the non-response is patterned.
Example: If at-risk adolescents are less likely to participate in an alcohol
consumption survey, the findings may underrepresent the actual alcohol
consumption rates among this group.

• Volunteer Bias: This type of bias happens when data points are self-
selected, as in voluntary surveys, skewing the dataset to a particular sub-
group with specific interests.
Example: If only adolescents concerned about their drinking habits vol-
unteer for a survey, the results may not accurately reflect the broader
adolescent population’s alcohol consumption behaviors.
• Sampling Frame Bias: This bias occurs when the sampling frame
doesn’t cover the entire target population.
Example: Conducting a survey only in schools misses adolescents who are
not enrolled, potentially biasing the results, as the school-going popula-
tion may differ in alcohol consumption habits from the general adolescent
population.
Each of these biases can significantly impact the accuracy of findings in
studies on adolescent alcohol consumption, making it essential to carefully
consider and address them in the data collection and analysis process.

References
[1] Syed Muhammad Sajjad Kabir et al. Basic guidelines for research. An
introductory approach for all disciplines, 4(2):168–180, 2016.

15
[2] Syeda Ayeman Mazhar, Rubi Anjum, Ammar Ibne Anwar, and Abdul Aziz
Khan. Methods of data collection: A fundamental tool of research. Journal
of Integrated Community Health, 10(1):6–10, 2021.
[3] Wendy Olsen. Data collection: Key debates and methods in social research.
2011.

[4] Hamed Taherdoost. Data collection methods and tools for research; a step-
by-step guide to choose data collection technique for academic and business
research projects. International Journal of Academic Research in Manage-
ment (IJARM), 10(1):10–38, 2021.

16

You might also like