Chapter 4
DATA COLLECTION
Comprehensive Study Notes
Research Techniques | Ashim Khadka
Topics Covered:
Primary & Secondary Data • Data Collection Methods • Data Preparation
Sampling Techniques • Evaluation & Validation of Data
1. Introduction to Data Collection
Data collection is the backbone of any research. Before we can analyze anything, draw
conclusions, or make decisions, we first need data. Without data, research is just guesswork.
💡 Analogy: Think of data collection like gathering ingredients before cooking. You cannot
bake a cake without flour, sugar, and eggs. Similarly, you cannot draw research conclusions
without first collecting the relevant data.
1.1 What is Data?
Data refers to raw facts, figures, or observations in their unprocessed form. On its own, data has no
meaning. It only becomes useful information once it is organized, processed, and interpreted.
✅ Example: The number '37' is just data. But '37 degrees Celsius body temperature of a
patient' becomes meaningful information.
Data can take many forms:
• Numerical: temperatures, scores, counts, percentages
• Textual: interview transcripts, survey responses, written records
• Visual: photographs, satellite images, scanned documents
• Auditory: recorded interviews, sound measurements
1.2 Types of Data
There are two fundamental types of data in research:
Feature Quantitative Data Qualitative Data
Definition Numerical values that can be Descriptive information that can be
measured or counted observed but not measured
Nature Structured, objective Unstructured, subjective
Collection Method Surveys, experiments, Interviews, observations, focus
measurements groups
Analysis Statistical analysis, graphs, charts Thematic analysis, coding,
interpretation
Examples Age: 25, Temperature: 36.6°C, Score: Customer satisfaction responses,
85% interview narratives
💡 Analogy: Think of quantitative data as a photo (precise, measurable) and qualitative
data as a painting (interpretive, rich in meaning). Both can capture reality but in different
ways.
2. Types of Data Collection: Primary vs Secondary
Once we know what type of data we need, the next question is: where do we get it from? There are
two main sources: Primary Data and Secondary Data.
Aspect Primary Data Secondary Data
Who Collects It? The researcher themselves Someone else (earlier research,
organizations)
Purpose Collected specifically for the current Originally collected for a different
research purpose
Freshness Up-to-date and directly relevant May be outdated or not perfectly
relevant
Cost Usually expensive and Generally cheaper, often freely
time-consuming available
Control Researcher has full control over No control; must rely on original
quality collector's accuracy
Examples Your own survey, interviews, lab Census data, government reports,
experiments published journals
2.1 Primary Data
Primary data is collected firsthand by the researcher specifically for the current research problem.
The researcher directly interacts with the source.
💡 Analogy: Primary data is like freshly cooked food — made exactly to your taste, fresh,
but takes time and effort to prepare.
Characteristics of Primary Data:
• Original: Never been published or used before; collected fresh for your specific purpose
• Relevant: Tailored precisely to your research questions
• Accurate: You control the quality of collection
• Expensive: Requires significant time, money, and human resources
• Time-consuming: You must conduct surveys, interviews, or experiments yourself
Methods of Collecting Primary Data:
• Interviews (structured, semi-structured, unstructured)
• Surveys and questionnaires
• Direct observations
• Experiments and lab studies
• Focus groups
✅ Example: A researcher studying the effect of social media on student grades designs
and distributes their own questionnaire to 200 students. The responses collected are primary
data.
2.2 Secondary Data
Secondary data is information that was originally collected by someone else for a different purpose.
The researcher accesses this existing data for their own research.
💡 Analogy: Secondary data is like reheated leftovers — convenient and fast, but may not
be exactly what you wanted, and you have no control over how it was originally prepared.
Sources of Secondary Data:
• Census Data: National population statistics, government records
• Academic Journals: Published research articles and papers
• Archival Records: Historical documents, old newspapers, official records
• Online Repositories: Datasets available on platforms like Kaggle, government open data
portals
• Reports: Company reports, NGO publications, WHO/UN documents
Advantages and Challenges:
• Advantage: Readily available without need for field collection
• Advantage: Cost-effective; saves time and resources
• Advantage: Can cover very long time spans (e.g., historical climate data)
• Challenge: Data may not perfectly match your current research objectives
• Challenge: Questions of currency — data may be outdated
• Challenge: Issues of accuracy — original data quality is unknown
✅ Example: A student researching Nepal's population growth uses the CBS (Central
Bureau of Statistics) census data from 2021. They did not collect it themselves — it is
secondary data.
3. Data Collection Methods
Now that we know WHERE data comes from (primary vs secondary), we need to understand HOW
to actually collect it. Different situations require different methods. There are 7 major data collection
methods.
📋 Overview: The 7 methods are: (1) Surveys, (2) Experiments, (3) Questionnaires, (4)
Interviews, (5) Observation, (6) Experiments, (7) Documents
3.1 Surveys
A survey is a systematic method of collecting data from a large group of people or events using
standardized questions or instruments. The goal is to collect enough data to identify patterns and
generalize findings to a broader population.
💡 Analogy: A survey is like casting a large net into the ocean. You cast it wide to catch
many fish (responses) at once, then study the patterns in your catch.
Key Features of a Survey:
• Involves a large number of respondents
• Uses standardized questions so everyone gets the same questions
• Results are analyzed statistically
• Findings can be generalized to a larger population
Planning and Designing Surveys — Six Key Steps:
1.Data Requirements
Define exactly what information you need. Data can be directly related to your research
question (e.g., opinions on online learning) or indirectly related (e.g., age, gender, education
level of respondents, which help interpret main results).
Important: In a survey, you normally get only one opportunity to collect data from respondents.
You cannot go back and ask again. So think ahead carefully about every question you need.
2.Data Generation Method
Choose how you will actually collect the data: questionnaires, interviews, observation, or
document review. Each has strengths and weaknesses (detailed below).
3.Sampling Frame
The sampling frame is the complete list of all members of your target population. From this
frame, you will select your sample.
✅ Example: If studying how an IT helpdesk handles queries, the sampling frame would be
the helpdesk's log of ALL requests ever submitted.
4.Sampling Technique
Decide HOW to select people from the sampling frame. There are many techniques (covered in
detail in Section 5).
5.Response Rate and Non-Responses
Not everyone will respond to your survey. Plan for this. If you need 100 responses, you may
need to send 200 invitations, accounting for non-responders.
6.Sample Size
Decide how many people to include. Larger samples give more reliable results, but are more
expensive. The size depends on your population size, desired accuracy, and available
resources.
3.2 Questionnaires
A questionnaire is a structured set of standardized written questions used to collect information from
respondents. It is usually self-administered (respondents fill it out themselves without a researcher
present).
💡 Analogy: A questionnaire is like a printed exam paper — everyone gets the same
questions in the same order, ensuring fair and consistent data collection.
Design Considerations for a Good Questionnaire:
• Clear Wording: Questions must be simple, unambiguous, and easy to understand by all
respondents
• Logical Order: Start with easy/general questions, then move to specific or sensitive ones
• Mix Question Types: Use a combination of closed questions (Yes/No, multiple choice) and
open questions (free text) to capture both quantitative and qualitative data
• Avoid Leading Questions: Don't phrase questions that push respondents toward a particular
answer
• Pilot Test First: Always test your questionnaire with a small group before full deployment
✅ Example: An HR manager distributes an employee satisfaction questionnaire via email.
It includes rating scales (1-5 for job satisfaction), multiple choice (department, years of
experience), and one open-ended question ("What would improve your work experience?").
3.3 Interviews
Interviews involve direct, personal interaction between the researcher and the respondent. Unlike
questionnaires, interviews allow the researcher to probe deeper, ask follow-up questions, and clarify
ambiguities.
💡 Analogy: If a questionnaire is like a multiple-choice exam, an interview is like an oral
viva — the examiner can dig deeper and explore nuanced answers.
Types of Interviews:
Type Description Flexibility Best Used For
Structured Fixed set of questions asked in Low — no Quantitative research;
exact order to all respondents deviation large samples needing
consistency
Semi-Structured Core questions planned, but Medium Exploring a topic while
interviewer can probe and explore maintaining some
based on responses consistency
Unstructured Conversational; no fixed questions; High — very Qualitative research;
flows naturally flexible exploring sensitive or
complex topics
✅ Example: A researcher studying shopping behavior interviews 20 customers using a
semi-structured interview. They have 5 planned questions but follow up with extra questions
based on interesting responses they hear.
3.4 Observation
In observation, the researcher collects data by directly watching behaviors, events, or processes
without actively intervening or manipulating them. The goal is to record what naturally happens.
💡 Analogy: Observation is like a nature documentary filmmaker. They set up cameras in
the jungle and quietly record animal behavior without disturbing it. They capture what really
happens, not what animals "say" they do.
Types of Observation:
• Participant Observation: Researcher joins the group being studied (e.g., researcher works
inside a call center to observe workflow)
• Non-Participant Observation: Researcher watches from outside without joining (e.g.,
watching customers from a distance in a store)
• Structured Observation: Uses a predefined checklist of behaviors to record
• Unstructured Observation: Open-ended; researcher records anything noteworthy
✅ Example: A researcher studying customer behavior in a supermarket stands near the
checkout counters and records how long customers wait, how they react to waiting, and
whether they abandon items. This is non-participant observation.
3.5 Experiments
An experiment is a controlled research method used to investigate cause-and-effect relationships.
The researcher deliberately manipulates one variable (the independent variable) and measures its
effect on another (the dependent variable).
💡 Analogy: An experiment is like a cooking test. You want to know if adding baking soda
makes bread rise higher. You bake two identical loaves — one with baking soda, one without
— and compare the results. Everything else stays the same. That's controlled
experimentation.
Key Concepts in Experiments:
• Independent Variable (IV): The variable you change or manipulate (e.g., amount of baking
soda)
• Dependent Variable (DV): The variable you measure to see the effect (e.g., height of bread)
• Control Group: Group not exposed to the change; serves as a baseline for comparison
• Experimental Group: Group exposed to the change being tested
• Hypothesis: A testable statement about the expected relationship between variables (e.g.,
"Adding baking soda will increase bread height")
6 Characteristics of Experiment-Based Research:
7.Observation and Measurement
Researchers make precise, detailed observations of outcomes when a specific factor is
introduced or removed. Everything is measured carefully.
8.Process of Experimentation
The process follows three stages:
◦ Measure/observe initial conditions (before)
◦ Manipulate circumstances (introduce the factor)
◦ Re-measure/re-observe to identify changes (after)
9.Proving or Disproving Relationships
The experiment aims to establish whether a causal relationship exists between the independent
and dependent variables.
10. Identification of Causal Factors
Identifies which variable (IV) causes change in another (DV), thus establishing cause and effect.
11. Explanation and Prediction
If an experiment supports the hypothesis, the established causal relationship can be used to
predict future outcomes.
12. Repetition
Experiments are repeated many times under varying conditions to ensure results are consistent
and not due to chance, faulty equipment, or external factors.
✅ Example: A researcher wants to test if a new teaching app improves students' math
scores. They divide students into two groups: Group A uses the app for 4 weeks
(experimental), Group B uses traditional methods (control). Both groups take the same math
test before and after. The difference in improvement is the effect of the app.
3.6 Documents
Document analysis (or document surveys) involves examining existing documents such as reports,
manuals, code repositories, meeting minutes, policies, contracts, and other written materials as a
source of data.
💡 Analogy: Using documents is like being a detective who studies old case files rather
than interviewing witnesses. You gather evidence from existing written records.
• Documents can be internal (company reports, internal memos) or external (published articles,
government legislation)
• Useful for understanding history, policies, official positions, and recorded events
• Sometimes researchers need to create new documentation (e.g., writing meeting notes during
an observed session)
✅ Example: A researcher studying how software companies manage bugs examines the
GitHub issue tracker logs and README files of 20 open-source projects to understand how
problems are documented and resolved.
4. Data Preparation
After collecting data, you cannot directly use it for analysis. Raw data is often messy, incomplete, or
inconsistent. Data preparation is the process of cleaning, transforming, and organizing raw data so
it is ready for accurate analysis.
💡 Analogy: Data preparation is like washing and chopping vegetables before cooking. You
wouldn't throw raw, muddy, uncut vegetables directly into the pot. You clean them, peel
them, and cut them to the right size first.
Why It Matters: Poor data preparation leads to inaccurate analysis results. The principle is
'Garbage In, Garbage Out' (GIGO) — if you feed bad data into analysis, you get bad
conclusions out.
4.1 Data Cleaning
Data cleaning involves finding and fixing problems in the raw dataset. This includes:
Problem Description Solution
Missing Values Some respondents skipped Impute (fill in estimates) or exclude the
questions or data wasn't incomplete records
recorded
Outliers Extreme values that are unusual Investigate — if clearly an error, correct or
(e.g., age = 150) remove it
Duplicate Records Same respondent or data point Remove duplicate entries
entered twice
Inconsistent Formats Same data recorded differently Standardize all values to one consistent
(e.g., 'Male', 'M', 'male', '1') format
Typos and Errors Spelling mistakes, wrong Correct errors using validation rules or
numbers, misclassifications manual review
✅ Example: A customer database has an entry where a customer's age is recorded as
150. This is clearly an error — it would be flagged and either corrected if the real age is
known, or excluded from age-related analyses.
4.2 Data Coding
Data coding means converting qualitative (descriptive) data into numerical codes so it can be
processed by statistical software.
Original Response Coded Value
Yes 1
No 0
Male 1
Female 2
Strongly Agree 5
Agree 4
Neutral 3
Disagree 2
Strongly Disagree 1
✅ Example: Survey Question: "Do you use online banking? Yes/No" — Coded as 1 (Yes)
and 0 (No). Now you can calculate that 78% of respondents (0.78 average) use online
banking.
4.3 Data Organization
After cleaning and coding, data must be organized in a format suitable for analysis software (Excel,
SPSS, R, Python, etc.). This involves:
• Structuring data in rows (individual records) and columns (variables)
• Ensuring consistent data types (numbers stored as numbers, not text)
• Creating a data dictionary that explains what each variable/code means
• Saving data in appropriate file formats (.csv, .xlsx, .sav)
Tip: Always pilot test your data collection instruments (questionnaire or interview guide) with
5-10 people before full deployment. This helps you discover data preparation problems
before they affect your entire dataset.
5. Data Sampling and Its Techniques
Even after defining your sampling frame (the total population), you usually cannot study EVERY
member of the population. Sampling is the process of selecting a manageable subset (sample) from
the larger population to study.
💡 Analogy: Sampling is like a doctor testing a patient's blood. The doctor doesn't drain all
of the patient's blood to analyze it — they take a small sample tube and analyze that. If the
sample is taken correctly, it accurately represents what's in the whole body.
Key Principle: A well-selected sample allows you to draw conclusions about the entire
population without studying every single member. This saves enormous amounts of time,
money, and resources.
5.1 Probability (Probabilistic) Sampling
In probability sampling, every member of the population has a known, non-zero chance of being
selected. This is the gold standard for research because it minimizes bias and allows statistical
inference to the full population.
5.1.1 Random Sampling
The simplest form. Each member of the population is equally likely to be selected. Selection is
entirely by chance, with no pattern or preference.
• Methods: Drawing names from a hat, using a random number generator, or random number
tables
• Strength: Unbiased; every member has an equal chance
• Weakness: Requires a complete list of all population members, which may be difficult to
obtain
✅ Example: A teacher wants to randomly select 10 students from a class of 50 to interview.
She writes all 50 names on slips of paper, puts them in a bowl, and draws 10. Each student
had an equal 10/50 = 20% chance of selection.
5.1.2 Systematic Sampling
Builds on random sampling by adding a regular interval. You select every k-th element from the
sampling frame, where k = population size ÷ sample size.
• Process: Randomly choose a starting point, then select every k-th element
• Strength: Simpler to execute than pure random sampling; evenly distributed across the
population
• Weakness: Can introduce bias if there's a hidden pattern in the list at the same interval
✅ Example: A researcher wants 100 people from a telephone directory of 10,000 entries. k
= 10,000/100 = 100. They randomly start at entry number 47, then select entries 147, 247,
347... every 100th entry. This is systematic sampling.
5.1.3 Stratified Sampling
The population is divided into distinct subgroups (strata) based on a characteristic (gender, age,
department). A random sample is then drawn from each stratum proportional to its size in the
population.
• Ensures every subgroup is represented in the sample
• Improves accuracy of estimates for each subgroup
• Best used when the population has distinct, meaningful subgroups
✅ Example: Studying employee stress in an IT company with 700 males (70%) and 300
females (30%). For a sample of 100 employees, you would randomly select 70 males from
the male list and 30 females from the female list. This ensures gender balance reflects the
real company.
💡 Analogy: Stratified sampling is like making a fruit salad. You don't just grab any fruits
randomly — you decide to include 3 apples, 2 oranges, and 1 mango to make sure all fruits
are represented fairly in proportion.
5.1.4 Cluster Sampling
The population is divided into naturally occurring groups called clusters (towns, schools,
neighborhoods, companies). A random selection of clusters is made, and then all or some members
within those clusters are studied.
• Clusters should ideally mirror the overall population
• Reduces travel and logistical costs significantly
• Risk: Bias if selected clusters are not representative of the whole population
✅ Example: A national study on student literacy would be too expensive if done school by
school across all 5,000 schools. Instead, researchers randomly select 50 schools (clusters)
and test all students in those schools. This is cluster sampling.
💡 Analogy: Cluster sampling is like randomly picking several bags of mixed candy instead
of picking individual candies one by one. You hope each bag is a good mix of the whole
range.
5.2 Non-Probability (Non-Probabilistic) Sampling
In non-probability sampling, not every member of the population has an equal chance of being
selected. Selection is based on researcher judgment, convenience, or other non-random factors.
These methods are faster and cheaper but introduce more bias and limit generalizability.
5.2.1 Purposive (Judgmental) Sampling
The researcher deliberately and intentionally selects participants who are most likely to provide
valuable, relevant, and information-rich data for the specific research purpose.
• Best for qualitative research where depth and specificity matter more than statistical
representativeness
• Requires researcher expertise to identify the 'right' participants
• Risk: Selection bias if criteria are not well-defined
✅ Example: A researcher studying online shopping behavior deliberately selects 10 people
who shop online every day AND 10 people who have never shopped online. These two
extreme groups will give rich, contrasting perspectives. This is purposive sampling.
5.2.2 Snowball Sampling
The researcher starts with one or a few participants who meet the study criteria. These initial
participants are then asked to recommend or refer others who also meet the criteria. The sample
grows like a rolling snowball.
• Ideal for studying hard-to-reach or hidden populations
• Very useful when there is no available sampling frame
• Risk: Participants may refer similar people (same social circles), reducing diversity
✅ Example: A researcher studying underground computer hackers cannot find them in a
directory. They find one hacker willing to participate and ask them to refer other hackers they
know. Those hackers refer more, and the sample snowballs from 1 to 25 participants.
💡 Analogy: Snowball sampling is like a chain referral. You get a job through a friend, who
heard about it from their colleague, who knew someone at the company. Each contact leads
to another.
5.2.3 Self-Selection Sampling
Participants voluntarily choose to participate in the study based on their own interest or motivation.
The researcher advertises the study and waits for volunteers.
• Participants opt in — they come to the researcher, not vice versa
• Useful when the researcher has limited access to the target population
• Risk: People who self-select often have strong opinions or special interest in the topic,
creating bias
✅ Example: A researcher posts a notice on a university noticeboard: "Volunteers needed
for a study on online learning experiences. Contact: researcher@[Link]." Students who
respond are self-selecting into the sample.
5.2.4 Convenience Sampling
The researcher selects participants who are most easily accessible and available at the time of the
study. No systematic or random selection is used.
• The easiest, fastest, and cheapest sampling method
• Appropriate for preliminary studies, pilot research, or initial hypothesis exploration
• Serious limitation: Sample may be highly unrepresentative of the broader population
✅ Example: A researcher studying mobile phone usage approaches 50 people at a local
shopping mall on a Tuesday afternoon and surveys them. These people are convenient to
access, but they may not represent all mobile phone users (e.g., working people are absent,
elderly may be underrepresented).
Method Selection Basis Bias Risk Best For
Random Pure chance Very Low Large-scale quantitative research
Systematic Regular intervals Low Ordered populations
Stratified Proportional Low Diverse populations with subgroups
subgroups
Cluster Natural groups Medium Geographically spread populations
Purposive Researcher Medium-High Qualitative, expert selection
judgment
Snowball Referrals High Hidden/hard-to-reach populations
Self-selection Volunteer opt-in High Limited-access populations
Convenience Easy availability Very High Pilot studies, quick surveys
6. Evaluation and Validation of Data
Collecting data is not enough. After collection and preparation, you must rigorously assess its
quality to ensure your research conclusions are trustworthy. This involves two related but distinct
processes: Evaluation and Validation.
Concept Definition Focus
Evaluation Assessing the overall quality of the data including Broad quality assessment of
accuracy, completeness, consistency, and the entire dataset
relevance
Validation Checking whether specific data values meet Specific checks on individual
predefined criteria or standards (formats, ranges, values or fields
cross-verification)
💡 Analogy: Think of it like quality control at a factory. Evaluation is the general quality
inspector who checks if the whole batch of products looks right. Validation is the precise
testing machine that checks each product meets exact specifications (correct weight,
dimensions, etc.).
6.1 Data Quality Evaluation — Three Pillars
6.1.1 Reliability
Reliability means consistency. If you collect data using the same method under the same conditions
at different times, you should get the same (or very similar) results. Unreliable data varies randomly
even when nothing real has changed.
✅ Example: A thermometer is reliable if it consistently reads 37°C every time you measure
the same room temperature. If it reads 35°C one minute and 40°C the next without the
temperature changing, it is unreliable.
• Test-retest reliability: Collect same data twice; compare results
• Inter-rater reliability: Two researchers observe same event; compare their records
6.1.2 Validity
Validity means accuracy — the data actually measures what it is supposed to measure. You can
have a reliable instrument that is consistently wrong (like a scale that always reads 2 kg heavier
than actual weight).
✅ Example: If you want to measure students' critical thinking skills but your test only asks
them to memorize facts, your instrument lacks validity — it's measuring memory, not critical
thinking.
• Check for errors such as typos, incorrect entries, or impossible values (e.g., age = 150)
• Internal validity: ensures changes observed are due to your intervention, not external factors
• External validity: ensures findings can be generalized beyond your specific sample
6.1.3 Completeness
Completeness means all required data fields are filled in. Missing data can severely compromise
your analysis, especially if the missing data is not random (e.g., older people tend to skip certain
questions).
✅ Example: In a health survey, if 30% of respondents skip the 'income' question, your
income-related analysis will be incomplete and potentially biased. You must decide: impute
missing values using statistical estimates, or exclude those records from income-related
analysis.
• Imputation: Filling missing values with estimates (e.g., mean of other responses)
• Exclusion: Removing records with missing critical fields
6.2 Validation Techniques
6.2.1 Cross-Validation
Cross-validation involves cross-checking your collected data against another independent data
source to confirm its accuracy.
✅ Example: You conduct an online survey about customer purchases. To validate the
results, you compare survey responses with actual sales records in the company database.
If survey results claim 80% bought Product A but sales records show only 50% did, there's a
discrepancy that needs investigation.
6.2.2 Pilot Testing
Before launching the full study, conduct a small-scale trial with a limited number of participants. This
tests whether your data collection instruments work as intended and reveals any problems in
advance.
✅ Example: Before sending a 30-question survey to 500 employees, you send it to 10
colleagues first. You discover that Question 12 is confusing and Question 20 has a
formatting error. You fix both before the main rollout. This is pilot testing.
• Identifies unclear or ambiguous questions
• Tests data entry processes
• Reveals technical issues in online surveys
• Estimates time required to complete
6.2.3 Triangulation
Triangulation means using multiple independent data collection methods to study the same
phenomenon. If different methods lead to the same conclusion, you can be more confident in your
findings.
✅ Example: A researcher studying customer satisfaction collects: (1) Questionnaire ratings
(quantitative), (2) In-depth interviews (qualitative), and (3) Mystery shopper observation
reports. All three sources agree that customers are frustrated with long wait times. This
three-way confirmation is triangulation.
💡 Analogy: Triangulation is like navigation at sea before GPS. Sailors used three
landmarks on land to triangulate their exact position. One landmark gives a general idea, two
narrow it down, but three give a precise fix. More methods = more certainty.
7. Chapter Summary — Quick Reference
Topic Key Points
Data Collection Backbone of research; data = raw facts; types: quantitative (numerical)
vs qualitative (descriptive)
Primary Data Firsthand collection; tailored to research; methods: surveys, interviews,
experiments; time-consuming and costly
Secondary Data Pre-existing data; census, journals, reports; cost-effective but may not
perfectly fit research
Surveys Systematic large-scale data collection; requires 6 steps: data
requirements, method, frame, technique, response rate, sample size
Questionnaires Standardized written questions; self-administered; needs clear wording
and logical order
Interviews Direct interaction; three types: structured, semi-structured,
unstructured; allows deep exploration
Observation Watch behaviors without intervention; participant or non-participant
Experiments Tests cause-effect; IV manipulated, DV measured; requires control
group; repeated for reliability
Documents Analyze existing written records; useful for historical and policy
research
Data Preparation Cleaning (remove errors), coding (convert to numbers), organizing
(structure for analysis)
Probability Sampling Equal chance for all: Random, Systematic, Stratified, Cluster
Non-Probability Sampling Unequal chance: Purposive, Snowball, Self-selection, Convenience
Evaluation Assess reliability (consistency), validity (accuracy), completeness (no
missing data)
Validation Techniques Cross-validation, Pilot testing, Triangulation
End of Chapter 4 Notes