Data Collection and Preprocessing
CONGRATULATIONS!
Dela Cruz, Jamica Ella D. 40/40
Abayata, Cyril James 37/40
Lazo, Queen Elizabeth T. 37/40
Yao, John Philip G. 37/40
Lalog, Kyle Andrew 35 /40
Navarro, Navelyn Audrey 35 /40
Ong, Xiaoxi D. 35 /40
Catibayan, Cameron James 34 /40
Topic outline
Data
Data sources Dealing with
Data quality and transformation
and collection missing values
cleaning and
methods and outliers
normalization
Data collection
the systematic process of gathering data or facts in an
organized manner to study, analyze, and draw
information from it.
It is the foundational step in research and
analysis.
• The accuracy, consistency, and integrity of data determine
the reliability of the results derived from it.
• The choice of data sources and collection methods
significantly influence the quality of data.
Write P for PRIMARY DATA and S for Secondary
Data
3.
1. direct observation or 2. published or unpublished
interview/questionnaires/rating 4. book
measurement materials
scales
5. newspapers 6. mail of recordings/videos 7. magazines 8. experimentation
10. birth/death/marriage
9. journals
certificates
Sources of data
Primary Data
data is collected directly from the source ,collected firsthand by the researcher for a specific research purpose or
objective.
It's raw data, meaning it hasn't been processed or analyzed yet.
tailored to the researcher's unique question
Secondary Data
Obtained from already existing sources.
may or may not fully align with the current research question
relevance and authenticity might be a concern, especially if the source is not credible
Data collection methods
Surveys/Questionnaires
structured with specific close-ended questions or unstructured with open-ended questions.
Delivery methods include online platforms, face-to-face interviews, telephonic surveys, or postal mails.
Interview
Direct, detailed, one-on-one conversations
Observations
collecting data by watching behavior or events. It can be participant (where the observer is part of the group) or non-participant.
Experiments
Often used in scientific studies, experiments involve a controlled setup where variables are manipulated to observe their effect on the subject.
Sampling
Collecting data from a subset of the population rather than the entire population.
Data collection methods
Case Studies
Detailed investigations of a single unit or case (e.g., an individual, community, or organization) to explore its complexities.
Focus Groups
moderated discussion with a group of individuals to gather diverse views on a topic.
Ethnographic Research
researcher studies the cultural phenomena by interacting closely with the community or group.
Digital and Web Analytics
collect data from digital platforms, like websites or apps, to analyze user behavior and other related metrics.
Content Analysis
Systematic analysis of the content of communications, like books, websites, and articles.
Identify the data collection methods used in the
following statements.
1. It is a procedure carried
3. is a research tool used to
out to support or refute a 2. It is a qualitative method
determine the presence of
hypothesis, or determine for collecting data often
certain words, themes, or
the efficacy or likelihood of used in the social and
concepts within some given
something previously behavioral sciences.
qualitative data (i.e. text).
untried.
4. It is a quantitative
5. A detailed study of a
measurements of the
specific subject, such as a
performance of online
person, group, place,
content, including
event, organization, or
advertising campaigns,
phenomenon.
social media, and websites.
Identify the data collection methods used in the
following statements.
6. A structured conversation where one participant asks questions, and the other provides
answers.
7. It is a method of gathering information using relevant questions from a sample of people
with the aim of understanding populations as a whole.
8. The act of watching something or someone carefully:
9. The act, process, or technique of selecting a suitable sample.
10. It is a small group of participants who share their feelings or perceptions on a particular
topic with researchers.
sampling
primary objective is to gather
process of selecting a subset of
relevant information without
individuals or items from a larger
having to investigate every Probability Sampling
population to draw conclusions
individual or item in the population,
about the entire population.
saving time and resources.
every member of the population Results obtained from a probability
has a probability of being selected. sample can often be generalized
Non-Probability Sampling
It provides the most accurate to the larger population with a
representation of the population. certain level of confidence.
not all members have a known or
equal chance of being included in
the sample. It's often used when
random sampling isn't feasible.
Probability Sampling or Non Probability Sampling
1. Every unit 2. Selection of 4. Simple
3. Convenience
has a chance of a sample from Random
Sampling
being selected. a population. Sampling
5. Cluster 6. Snowball 7. Systematic 8. Quota
Sampling Sampling sampling Sampling
9. Stratified
10. Purposive
Random
Sampling
Sampling
Proba and
non-proba sampling PROBABILTY SAMPLING
techniques Simple Random Sampling- Every individual/item has an equal chance of being
selected.
NON-PROBABILITY SAMPLING
It's like a random draw or lottery. Tools such as random number generators are
1. Convenience Sampling- Selecting the most
often used.
accessible individuals/items. It's quick but not very
reliable.
2. Stratified Random Sampling- The population is divided into mutually exclusive
subgroups (strata) and then random samples are drawn from each stratum.
2. Judgmental or Purposive Sampling- The
researcher selects specific individuals/items based on
their judgment. It's often used in qualitative research This method ensures that each subgroup is adequately represented.
where specific perspectives are sought.
3. Snowball Sampling- Used often for hard-to-reach 3. Cluster Sampling- The population is divided into clusters, often geographically.
populations. An initial participant refers other
participants, who then refer more, creating a "snowball"
effect. Then, a random sample of clusters is chosen, and all individuals/items within
those chosen clusters are surveyed.
4. Quota Sampling- The researcher ensures equal or
proportionate representation of subjects based on 4. Systematic Sampling- Selecting every nth individual/item from a list.
certain characteristics, similar to stratified sampling, but
without random sampling within the strata.
For instance, every 10th person on a list might be selected.
DETERMINING SAMPLE SIZE
to determine the sample size
required to detect a statistically
incorporate the desired level of
significant effect if it exists. It takes
confidence, margin of error,
1. Statistical Formulas 2. Power Analysis into account factors such as the
population variability, and other
effect size, significance level
relevant parameters.
(alpha), and statistical power (1 -
beta).
experimental methodologies
require at least 15
similar previous studies or
participants according to Cohen et
conducting pilot studies can
3. Previous Studies or Pilot al. (2007:102), and there should be
provide insights into the variability 4. Rules of Thumb
Studies at least 15 participants in control
and effect sizes expected in your
and experimental groups for
research.
comparison according to Gall et al.
(1996).
According to Fraenkel Wallen, a
correlational study's minimum
acceptable sample size is at least
30.
A-SF, B-PA, C-PS, D-RT
1. Slovin’s Formula - n=N/1+(Ne2)
2. It is the calculation used to estimate the smallest sample size needed for an
experiment, given a required significance level, statistical power, and effect size. It helps
to determine if a result from an experiment or survey is due to chance, or if it is genuine
and significant.
A-SF, B-PA, C-PS, D-RT
3. Refers to a collection of research and studies that dealt
with the issue that the researcher investigated, and these
studies give the researcher with a wealth of information on
the subject of the study that aids him in fully comprehending
the subject of his scientific research.
4. It is an approximate method for doing something, based on
practical experience rather than theory.
Data Quality
1. Accurate: The data correctly represents the real-world
scenario.
2. Complete: There aren't missing values or parts of data’
3. Relevant: The data is appropriate for the task at hand.
4. Timely: The data is up-to-date.
5. Consistent: There aren't contradictions within the data.
Data Preprocessing
• before any formal analysis can be performed on the data, it often
needs
to be cleaned, transformed, and otherwise prepared.
• involves removing errors, handling missing values, or changing the
format
of the data.
• the techniques used to convert raw data into a clean data set.
Data Preprocessing
Importance of Data Preprocessing
1. Improving Data Quality-Raw data often comes with noise, errors, or
inconsistencies. Preprocessing can help address these issues, ensuring that the
data you're analyzing or modeling is of high quality.
2. Enhancing Efficiency- Algorithms and models perform better and faster when
working with preprocessed data. This is particularly crucial for large datasets or when
deploying models in real-time applications.
3. Achieving Better Results- High-quality data leads to more accurate models and
insights. If your data is not preprocessed correctly, even the best algorithms might
not produce meaningful results.
4. Facilitating Data Integration- combining data from different sources, preprocessing
can help harmonize different data formats, structures, or standards.
5. Reducing Model Complexity: Feature engineering, a subset of preprocessing, can
reduce the dimensionality of the data. This not only improves performance but can
also lead to simpler models that are easier to interpret.
Steps for data preprocessing
looking at your data closely to removing unnecessary and adding more useful data or
understand its overall "quality". redundant data to make improving the existing data to
How much data is missing? analysis manageable and make it more valuable for
What kind of values are there? faster your analysis.
cleaning the data by fixing errors changing the format, checking to ensure the
and filling in or removing any structure, or values of data is accurate and
missing values. data to make it suitable suitable for your
for analysis analysis.
Data preprocessing techniques
Data cleansing- Techniques for cleaning up FEATURE ENGINEERING- Extracting useful features from given data
messy data • CREATING NEW FEATURES- EX. : If you have date data like
"purchase date" and "delivery date", you might create a new feature
• Identify and sort out missing data called "delivery time" by subtracting the two.
• fixing outliers • FEATURE INTERACTION- EX. : Combining two features (like
• Reduce noisy data-meaningless or corrupt "length" and "width" to create "area") gives more valuable
data information.
• SCALING AND NORMALIZING-ENSURING ALL FEATURES
• Identify and remove duplicates. HAVE A CONSISTENT SCALE.
• DIMENSIONALITY REDUCTION-REDUCE DIMENSIONS TO
CAPTURE THE ESSENCE AND MAKE THE DATA MORE
MANAGEABLE.
• ENCODING-EX: DATA, LIKE CATEGORIES (E.G., "MALE" AND
"FEMALE"), NEED TO BE TRANSFORMED INTO NUMBERS
(E.G., 0 AND 1) SO THAT ALGORITHMS CAN UNDERSTAND AND
UTILIZE THEM.
Missing Values
WHY IS DATA MISSING?
• values or data that is not • PAST DATA MIGHT GET CORRUPTED
• DUE TO IMPROPER MAINTENANCE.
stored (or not present) for • OBSERVATIONS ARE NOT RECORDED FOR
some variable/s in the CERTAIN FIELDS DUE TO SOME REASONS.
THERE MIGHT BE A FAILURE IN
given dataset. RECORDING THE VALUES DUE TO
HUMAN ERROR.
• THE USER HAS NOT PROVIDED THE
VALUES INTENTIONALLY
• ITEM NONRESPONSE: THIS MEANS THE
PARTICIPANT REFUSED TO RESPOND.
Handling Missing Values
Deleting the Missing value
Deleting the entire row (listwise deletion)
Deleting the entire column
Imputing the Missing Value
Replacing with an arbitrary value
Replacing with the mean , median (numeric) or mode
(categorical)
Replacing with the previous value -mostly used in time
series
Identifying Outliers
Box Plots and Scatter Plots- Visualize data, use box plots and scatter plots to
identify data points that fall far from the median or mean.
Z-Score Analysis- Calculate the z-score for each data point to measure how many
standard deviations it is away from the mean. Set a threshold to identify potential
outliers.
IQR (Interquartile Range) Method- Calculate the IQR and identify data points
outside the acceptable range (typically defined as 1.5 times the IQR).
Handling Outliers
Removing Outliers- consider removing them from the dataset. However, be cautious
not to remove too many data points.
Transformations- Apply mathematical transformations like logarithm, square root, or
Box-Cox to bring extreme values closer to the mean. This can help normalize the data.
Capping and Flooring- Set limits for the acceptable range of values. Any data point
outside this range can be capped or floored at the respective limit.
Mean/Median Imputation- Replace outlier values with the mean or median of the
column
Imputation Based on Predictive Models-Use regression or other predictive models to
estimate values for outliers based on other relevant variables in the dataset.
Multiple Imputation: Generate multiple imputed datasets by creating multiple
estimates for missing values.
Data Transformation
process of changing the format, structure, or values of data, making
data more suitable for analysis, making it compatible with certain tools,
or preparing it for modeling.
Standardization - transforms features to have a mean (average) of 0
and a standard deviation of [Link] to a Range (Min-Max Scaling) -
adjusts features so that they fall within a given range, usually [0,1].
Log Transformation - Useful for dealing with data that spreads out in
increasing intervals (exponential growth)
Normalization is a subset of data transformation, specifically focusing
on scaling numeric data to a standard format or scale.
Always document the decisions you make during the
data cleansing process to ensure transparency and
reproducibility.