Chapter 8: Analyzing Quantitative Data —
Quantitative research deals with:
Measurements
Numerical values
Statistical interpretation
Researchers use statistics to:
Summarize data
Discover patterns
Make predictions
Compare results
Determine significance of findings
The chapter emphasizes that:
“Numbers are meaningless unless we can find the patterns beneath them.”
Statistics help researchers answer:
“What do the data mean?”
EXPLORING AND ORGANIZING A
DATA SET
Before doing any statistical calculations, researchers should:
Observe data carefully
Think critically
Look for hidden patterns
Organize data meaningfully
The chapter strongly stresses that:
Careful observation of data is extremely important.
“How researchers organize data affects the meaning
revealed by those data.”
Time-series study
Meaning:
Data are collected repeatedly over time.
Such studies often reveal clear visual patterns.
USING COMPUTER SPREADSHEETS
Organizing large amounts of data manually is:
Difficult
Slow
Time-consuming
Modern researchers use spreadsheet software like:
Microsoft Excel
LibreOffice
Spread32
Advantages of Electronic Spreadsheets
Once data are entered:
Calculations become automatic
Changes update automatically
Data can be reorganized easily
Major Spreadsheet Functions
1. Sorting
Data can be rearranged:
Highest to lowest
Alphabetically
By age
By category etc.
Example:
Math scores can be sorted from youngest to oldest students.
2. Recoding
Researchers can transform old data into new categories.
Example:
Age Group
7–9 Group 1
10–12 Group 2
13–15 Group 3
3. Formulas
Spreadsheets can calculate:
Mean
Totals
Percentages
Standard deviation
Many formulas are already built into the software.
4. Graphing
Software automatically creates:
Line graphs
Pie charts
Bar graphs
Researchers can customize:
Axes
Labels
Legends
5. “What If” Analysis
Researchers can test multiple possibilities quickly.
Examples:
Compare males vs females
Compare different age groups
Remove one subgroup and reanalyze
This makes experimentation easier.
Researchers must choose statistics according to:
Nature of data
Research purpose
Scale of measurement
Distribution pattern
Looking at data in only one way gives an
incomplete understanding.
Therefore researchers use multiple statistical techniques.
FUNCTIONS OF STATISTICS
Statistics have TWO major functions:
1. Descriptive Statistics
These describe and summarize data.
They explain:
Center of data
Spread of data
Relationships among variables
Examples:
Mean
Median
Mode
Standard deviation
2. Inferential Statistics
These help researchers:
Draw conclusions about populations
Use sample data to make predictions
Test hypotheses
Researchers study small samples and generalize findings to larger populations.
Example of Inferential Statistics
An immigration officer meets only some Egyptians.
From this small sample, the officer forms ideas about Egyptians in general.
Similarly:
Researchers study samples to infer characteristics of entire populations.
Statistics as Estimates of Population
Parameters
Parameter
A parameter is:
A characteristic of an entire population.
Examples:
Population mean
Population standard deviation
Statistic
A statistic is:
A value calculated from a sample.
Researchers use sample statistics to estimate population parameters.
Statistical Symbols
Characteristic Population Symbol Sample Symbol
Mean μ M or X̄
Standard deviation σ s or SD
Proportion P p
Total number N n
Single-Group Data
A study produces single-group data when information is collected from only one group.
Example
Test scores of one classroom
Heights of students in one school
Blood pressure readings of one group of patients
In such studies, researchers usually analyze:
Average performance
Variability
Trends within the same group
Multi-Group Data
A study produces multi-group data when researchers compare two or more groups.
Example
Boys vs girls
Experimental group vs control group
Private school students vs government school students
Researchers mainly examine:
Differences between groups
Similarities between groups
Effects of treatments or conditions
Different statistical techniques are required for single-group and multi-group data.
Because:
Single-group studies focus on describing one group.
Multi-group studies focus on comparing groups.
Continuous vs Discrete Variables
A variable is any characteristic that can change or vary.
Example:
Age
Height
Number of students
Marks
Variables are mainly of two types:
1. Continuous Variable
A continuous variable can take infinite possible values within a range.
You can measure it very precisely.
Infinite possible values
Values can lie between two numbers
Can contain decimals/fractions
Examples
Age
A person can be:
18 years
18.5 years
18.75 years
18 years 2 months etc.
So age is continuous.
Height
Examples:
160 cm
160.2 cm
160.25 cm
Infinite possibilities exist between two heights.
Weight,temp
2. Discrete Variable
A discrete variable has only specific separate values.
No values exist between them.
Countable values
No fractions usually
Separate distinct categories
Usually counted, not measured
Examples
Number of Students
Possible values:
20 students
21 students
22 students
Not possible:
20.5 students
Grade Levels
9th
10th
11th
12th
No “9.5th class.”
Nominal, Ordinal, Interval, and Ratio
Scales
These are scales of measurement.
They determine:
Type of data
Statistical methods usable
Mathematical operations possible
Nominal Scale
Numbers represent categories only.
Examples:
Male = 1
Female = 2
Here, 2 is not greater than 1. The numbers are only labels.
Other examples:
Religion
Political affiliation
Blood group
Suitable statistic
Mode
2. Ordinal Scale
Ordinal data show order or ranking.
Example:
Class rank
Competition positions
If a student has Rank 1 and another has Rank 2, we know the first student performed better.
However, we do not know how much better
Suitable statistics
Median
Percentiles
3. Interval Scale
Interval data have equal intervals between values.
Example:
IQ scores
Temperature in Celsius
Difference between:
10°C and 20°C
is equal to
20°C and 30°C
But zero is not a true zero.
0°C does not mean absence of temperature.
Features
Order exists
Equal intervals exist
No true zero point
Suitable statistics
Mean
Standard deviation
Correlation
4. Ratio Scale
Ratio data are similar to interval data but also have a true zero point.
Example:
Height
Weight
Income
If income = ₹0, it means no income.
Ratios are meaningful:
40 kg is twice 20 kg.
Features
Order exists
Equal intervals exist
True zero exists
Most advanced scale
Suitable statistics
All statistical techniques can be used.
Normal Distribution (Bell Curve)
Many natural and human characteristics follow this pattern.
Examples:
Height
IQ
Crop production
Test scores
Biological traits
The graph looks bell-shaped and symmetrical.
Characteristics of Normal Distribution
1. Symmetrical Shape
Left side mirrors right side.
Both halves are equal.
This means:
Data are evenly distributed around the center.
2. Highest Point at Center
Most values occur near the middle.
Very high or very low values occur less frequently.
Thus:
Average values are most common.
3. Mean = Median = Mode
In a perfect normal curve:
Mean
Median
Mode
all occur at the same central point.
Percentages in Normal Distribution
The chapter explains how population percentages are distributed around the mean using
standard deviations.
Standard deviation helps determine:
How spread out data are
How close scores are to the average
Smaller SD:
→ Scores tightly clustered
Larger SD:
→ Scores widely spread
Between Mean and 1 Standard Deviation
About:
34.1%
lies on EACH side.
So together:
68.2%
of data lies within 1 standard deviation of the mean.
Between 1 SD and 2 SD
About:
13.6%
lies on each side.
Beyond 2 SD
Only:
2.3%
lies at each extreme end.
This means:
Extremely high or low values are rare.
NON-NORMAL DISTRIBUTIONS
Not all data follow a perfect bell curve.
Some distributions are distorted.
SKEWED DISTRIBUTIONS
Skewness means:
One side of the distribution stretches farther than the other.
The stretched side is called: Tail
Positive Skew
In positive skew:
Tail extends toward the right
Peak lies toward left
Usually caused by:
A few extremely high scores
Example:
Income distribution
Most people earn average salaries.
Few earn extremely high salaries.
Negative Skew
In negative skew:
Tail extends toward left
Peak lies toward right
Caused by:
A few extremely low scores
KURTOSIS
Kurtosis refers to:
How peaked or flat a distribution is.
Leptokurtic Distribution
Very peaked
Tall center
Data tightly clustered
Meaning:
Most values are close to mean.
Platykurtic Distribution
Flat and broad
Data widely spread
Meaning:
Greater variability exists.
PERCENTILE RANKS
Percentile ranks are often used in:
Achievement tests
Aptitude tests
Meaning of Percentile Rank
A percentile tells:
How many people scored below a person.
Example:
90th percentile means:
Person performed better than 90% of people.
Formula for Percentile Rank
Percentile ranks are ordinal data.
PARAMETRIC VS NONPARAMETRIC
STATISTICS
Statistical methods are divided into two categories:
1. Parametric statistics
2. Nonparametric statistics
PARAMETRIC STATISTICS
Parametric statistics assume certain things about data.
Two major assumptions:
Assumption 1
Data should use:
Interval scale OR
Ratio scale
Assumption 2
Data should follow:
Normal distribution
If assumptions are violated:
Results may become inaccurate.
Examples of Parametric Statistics
t-test
ANOVA
Pearson correlation
NONPARAMETRIC STATISTICS
Nonparametric statistics:
Do NOT require normal distribution
Can handle ordinal data
Can handle skewed distributions
Useful when assumptions fail.
Why Not Use Nonparametric All the Time?
Because:
Parametric statistics are more powerful and sophisticated.
DESCRIPTIVE STATISTICS
THREE MAJOR MEASURES
1. Measures of central tendency
2. Measures of variability
3. Correlation measures
MEASURES OF CENTRAL TENDENCY
Central tendency means:
The center point around which data cluster.
It represents the “typical” value.
Three Main Measures
1. Mode
2. Median
3. Mean
Each has different uses.
MODE
Definition:
Most frequently occurring value.
Example
Data:
3, 4, 6, 7, 7, 9, 9, 9, 9, 10, 11
Mode = 9
because:
9 appears most often.
Features of Mode
Advantages
Simple
Useful for nominal data
Disadvantages
May not represent true center
Unstable across samples
MEDIAN
Definition:
Middle value in ordered data.
Finding Median
Arrange data:
Lowest to highest
Then:
Find middle score
If Number of Scores is Odd
Median =
Exact middle score
If Number of Scores is Even
Median =
Average of two middle scores
Uses of Median
Median is best for:
Ordinal data
Highly skewed data
Why Median Works Better in Skewed Data
Extreme values affect mean heavily.
Median ignores extremes.
Example:
3, 4, 5, 5, 6, 9, 15, 17, 125
Mean of Above Data
Mean becomes:
21
But most scores are actually near:
5–9
So mean becomes misleading.
Median of Same Data
Median = 6
This better represents center.
MEAN (Arithmetic Average)
Definition:
Arithmetic average of all scores.
Formula for Mean
Important Feature of Mean
Mean uses:
Every score in the dataset.
Limitation of Mean
Sensitive to:
Extreme scores (outliers)
GEOMETRIC MEAN
Used for:
Growth phenomena
Examples:
Population growth
Biological growth
Compound interest
Formula for Geometric Mean
Growth Curves
Stages:
1. Slow beginning
2. Rapid acceleration
3. Leveling off
Researchers must choose statistics according to:
Data characteristics
Distribution shape
Measurement scale
NOT according to personal preference.
MEASURES OF VARIABILITY
Variability means:
How spread out the data are.
It tells us:
Whether scores are close together
OR
Widely scattered
Two datasets may have the same mean but very different variability.
Importance of Variability
Measures of central tendency alone are insufficient.
Example:
Two classes may both have:
Mean = 75
But:
Class A
Scores:
74, 75, 76
Class B
Scores:
40, 60, 75, 90, 110
Same mean, but very different spread.
Thus:
Researchers must measure
dispersion/spread also.
THE RANGE
The simplest variability measure is:
Range
Definition:
Difference between highest and lowest score.
Formula for Range
Example Showing Problem with Range
Family sizes:
1, 3, 3, 3, 4, 4, 5, 5, 6, 15
Range:
15−1=1415-1=1415−1=14
But most families actually have:
3–6 children.
The extreme value 15 distorts interpretation.
INTERQUARTILE RANGE (IQR)
To avoid extreme-score problems, researchers use:
Interquartile Range
Quartiles
Data are divided into:
Four equal parts
Quartile 1 (Q1)
25% of scores lie below it.
Quartile 2 (Q2)
Middle point.
Equivalent to:
Median
50% above and 50% below.
Quartile 3 (Q3)
75% of scores lie below it.
Formula for IQR
Meaning of IQR
Shows spread of:
Middle 50% of data
Average Deviation
Represents:
Average distance of scores from mean
STANDARD DEVIATION
The most important variability measure.
Definition:
Average spread of scores around the mean.
Variance
Variance =
Average of squared deviations
Standard Deviation
Standard deviation =
Square root of variance
This returns the value to original measurement units.
Small SD:
→ Scores clustered closely
Large SD:
→ Scores spread widely
Importance of Standard Deviation
Used in:
Inferential statistics
Hypothesis testing
Normal distribution
Correlation
Regression
It is one of the most important statistics.
VISUAL INTERPRETATION OF
VARIABILITY
Low Variability
When scores cluster tightly around mean:
Data are more similar
More homogeneous
High Variability
When scores spread widely:
Data become more diverse
Less representative of average
Z SCORES (STANDARD SCORES)
A z score tells:
How far a score lies from the mean in SD
units.
Formula for z Score
Where:
X = individual score
M = mean
s = standard deviation
Positive z Score
Means:
Score is above mean
Example:
z = +2
→ score is 2 SD above mean.
Negative z Score
Means:
Score below mean
Example:
z = -1
→ score is 1 SD below mean.
z=0
Means:
Score exactly equals mean.
IQ Score
The book gives IQ scores as an example of:
Interval Scale data
IQ (Intelligence Quotient) scores are assumed to have:
Equal intervals between values.
Example:
IQ Scores
85
95
105
115
The difference between each pair is 10, so the intervals are equal.
This means:
Difference between 85 and 95
is considered equal to
Difference between 105 and 115
Important Point About IQ Scores
The book explains that:
Zero is not a true zero in IQ scores.
Stanine
Stanine means:
“Standard Nine”
It is a method of reporting test scores by dividing performance into 9 categories.
Stanines help simplify interpretation of scores.
Features of Stanine
Scores range from 1 to 9
Middle stanine = 5
Based on normal distribution
Used in educational and psychological testing
CORRELATION
Correlation studies relationships between variables.
“When one variable changes, does another
change too?”
Examples
Height and weight
Study time and marks
Temperature and ice cream sales
Positive Correlation
As one variable increases:
Other also increases.
Example:
Study time ↑
Marks ↑
Negative Correlation
As one variable increases:
Other decreases.
Example:
Stress ↑
Sleep ↓
No Correlation
Variables unrelated.
Change in one does not affect other.
Correlation Coefficient
Represented by:
r
Range:
Interpretation of r
Value Meaning
+1 Perfect positive correlation
0 No correlation
-1 Perfect negative correlation
Important Limitation of Correlation
The book gives an extremely important warning:
Correlation statistics are based on the
assumption that the relationship between
variables is linear.
Linear Relationship
A linear relationship means:
Relationship forms a straight line.
As one variable changes:
Another changes at constant rate.
Problem with Nonlinear Relationships
Some relationships are:
Curved instead of straight.
In such cases:
Correlation may fail to detect the relationship properly.
Example from Book: BMI and Anxiety
The chapter gives an example using:
Body Mass Index (BMI)
Anxiety level
Researchers found:
Anxiety was highest in people who were:
o Very underweight
o Very overweight
Anxiety was lowest in people with:
o Average BMI
This forms:
U-shaped relationship
not a straight-line relationship.
Why Correlation Fails Here
A regular correlation coefficient assumes:
Linear relationship
But BMI and anxiety show:
Nonlinear relationship
Therefore ordinary correlation may not properly represent the association.
Scatter Plots
Correlations are often shown using:
Scatter diagrams
Each point represents:
One participant’s scores on two variables.
Strong Positive Correlation
Points form upward pattern.
Strong Negative Correlation
Points form downward pattern.
Weak Correlation
Points scattered randomly.
IMPORTANT WARNING
The chapter emphasizes:
Correlation does NOT prove causation.
Just because two variables relate does not mean one causes the other.
Example
Ice cream sales and drowning deaths may both increase in summer.
But:
Ice cream does not cause drowning.
Summer temperature affects both.
Coefficient of Determination
The chapter introduces:
Coefficient of determination
Represented by:
r²
how much one variable can be explained or
predicted by another variable.
Example
Meaning:
64% variability is shared.
Larger r2→ stronger relationship
Smaller r2→ weaker relationship
INFERENTIAL STATISTICS
Inferential statistics allow researchers to:
Draw conclusions about large populations
Use only small samples for analysis
Estimate population characteristics
Test hypotheses statistically
Two Main Functions
1. Estimating population parameters
2. Testing statistically based hypotheses
ESTIMATING POPULATION
PARAMETERS
Researchers usually study only a sample, not the whole population.
Example: Connecting-Rod Pins
A company manufactures metal pins for machinery.
Problem:
Some pins may be faulty.
Researcher Jan selects a random sample.
She wants to estimate:
1. Average pin diameter
2. Variation in diameter
3. Proportion of acceptable pins
This is estimation of population parameters using sample data.
Important Assumption
Inferential statistics assume that:
The sample is randomly selected and representative of the population.
SAMPLING DISTRIBUTION OF THE
MEAN
When many random samples are taken:
each sample has its own mean
those sample means form a distribution
This is called the sampling distribution of the mean.
Key idea:
Even if individual scores vary a lot,
sample means vary less.
STANDARD ERROR OF THE MEAN
The standard deviation for the distribution of sample means is called
Standard Error of the Mean (σM)
Formula:
Where:
σ = population standard deviation
N = sample size
Meaning of Standard Error
Small standard error → sample mean close to population mean
Large standard error → less accurate estimation
Increasing sample size decreases standard error.
Estimated Standard Error
Usually we do not know the population standard deviation
So we estimate it using the sample standard deviation
The formula becomes:
Where:
S = sample standard deviation
N = sample size
POINT ESTIMATES vs INTERVAL
ESTIMATES
A. Point Estimate
Single value used to estimate population parameter
example:
Sample average diameter = 0.712
So Jan says:
“I think the population mean is 0.712.”
This ONE number is called:
Point Estimate
This becomes estimate of population mean.
Problem:
Exact estimate is unlikely to be perfectly accurate.
B. Interval Estimate
Gives a range within which population parameter probably lies.
Example:
“The true population mean is probably between 0.708 and 0.716.”
Confidence Interval
This interval usually has a confidence level.
Example:
95% Confidence Interval
Means:
Researcher is 95% confident that the true population mean lies inside this range.
Example:
0.708 to 0.716
Understanding Confidence Intervals
If sample means form normal distribution:
About:
68% lie within 1 standard error
95% lie within 2 standard errors
Thus researcher can estimate likely population range.
HYPOTHESIS TESTING
Second major function of inferential statistics.
TWO TYPES OF HYPOTHESES
1. Research Hypothesis
A prediction or educated guess made by researcher.
Example:
“Teaching method A improves scores.”
Purpose:
Guides research
2. Statistical Hypothesis
Used during statistical testing.
Usually a:
Null Hypothesis (H₀) there is no effect, no difference, or no
relationship between variables.
Null hypothesis states:
Any observed difference occurred only due to chance.
Example:
Two group means differ only because of random variation.
TESTING THE NULL HYPOTHESIS
Researchers compare:
actual observed results
with
results expected by chance
If probability of occurring by chance is very small,
the null hypothesis is rejected.
Statistical Significance
Researchers calculate probability that results happened by chance.
If probability is very small:
Result is considered statistically significant.
Significance Level (Alpha, α)
Common alpha levels:
Alpha Level Meaning
.05 5% chance result occurred randomly
.01 1% chance
.001 0.1% chance
If:
p < .05
→ Reject null hypothesis.
Example
If probability of result due to chance is:
1 in 1000
then:
likely NOT due to chance
some real factor caused the difference
TYPE I AND TYPE II ERRORS
Type I Error (Alpha Error)
we believe
there is realRejecting null hypothesis when it is actually true.
effect or diff
Meaning:
Concluding effect exists when it really does not.
Example:
Researcher says medication works when it actually doesn’t.
False postive
Probability of Type I error = α
Type II Error (Beta Error)
Failing to reject null hypothesis when it is false.
Meaning:
Concluding no effect exists when it actually does.
False negative
Example:
Medication actually works, but researcher concludes it doesn’t.
Trade-off Between Errors
Reducing one error often increases the other.
Lower α:
fewer Type I errors
more Type II errors
Higher α:
more Type I errors
fewer Type II errors
POWER OF A STATISTICAL TEST
Power = ability to correctly reject false null hypothesis
If something REAL is happening,
a powerful test can detect it..
Ways to increase power:
1. Use Larger Sample Size
Larger samples:
Reduce sampling error
Produce more accurate estimates
2. Improve Validity and Reliability
Better measurements:
Reduce error
Increase likelihood of detecting true effects
3. Use Parametric Statistics When Possible
Parametric tests are generally more powerful.
Nonparametric tests:
Require fewer assumptions
But are less powerful
.
RESEARCH HYPOTHESIS vs NULL
HYPOTHESIS
Often they are opposites.
Example:
Research hypothesis:
Groups are different
Null hypothesis:
Groups are the same
Researchers indirectly support research hypothesis by rejecting null hypothesis.
IMPORTANT POINT
Statistics alone are not enough.
Researchers must:
interpret findings
explain meanings
connect results to research problem
EXAMPLES OF INFERENTIAL
STATISTICAL PROCEDURES
PARAMETRIC STATISTICS
Used for interval/ratio data and normal distributions.
1. Student’s t-test
Purpose:
Compare two means
Types:
Independent samples t-test
Dependent samples t-test
2. ANOVA (Analysis of Variance)
Purpose:
Compare 3 or more means
Uses F-ratio.
If significant:
follow-up post hoc tests needed
3. ANCOVA
Analysis of covariance.
Purpose:
Compare means while controlling another variable (covariate)
More statistically powerful than ANOVA.
4. Regression
Predicts dependent variable using independent variable(s).
Types:
Simple regression
Multiple regression
Important:
Correlation ≠ causation.
5. Factor Analysis
Finds clusters of related variables called factors..
Used to identify underlying factors/themes.
7. Structural Equation Modeling (SEM)
Advanced statistical technique.
Examines:
Complex relationships among variables
Direct and indirect effects
Can include:
Mediating variables
Moderating variables
NONPARAMETRIC STATISTICS
Used when:
assumptions for parametric tests are violated
data are ordinal or nominal
1. Mann–Whitney U Test
Compares medians of two groups.
Nonparametric equivalent of independent t-test.
2. Kruskal–Wallis Test
Compares medians of 3+ groups.
Equivalent of ANOVA.
3. Wilcoxon Signed-Rank Test
Compares correlated ordinal variables.
Equivalent of dependent t-test.
4. Chi-Square Test (χ²)
Tests how closely observed frequencies match expected frequencies.
Applicable for:
nominal
ordinal
interval
ratio data
5. Odds Ratio
Measures relationship between two dichotomous variables.
Example:
smoking vs heart disease
6. Fisher’s Exact Test
Used for:
very small samples (n < 30)
Alternative to chi-square/t-test for small samples.
META-ANALYSIS
Meta-analysis = analysis of analyses.
Researchers combine results from many previous studies.
Useful when:
many studies already exist on same topic.
STEPS IN META-ANALYSIS
1. Search for Relevant Studies
Systematic search using:
journals
databases
keywords
2. Select Appropriate Studies
Include only studies matching:
treatment
variables
population
methods
3. Convert Results to Common Statistical
Index
Different studies may use:
t-tests
ANOVA
regression
Researcher converts findings into common index.
Usually:
Effect Size (ES)
Effect Size
Effect size tells:
How strong or large the effect is.
Example
Medicine may:
Reduce stress slightly
Reduce stress greatly
Effect size measures this magnitude.
IMPORTANCE OF META-ANALYSIS
Helps researchers:
summarize large bodies of research
identify overall trends
resolve conflicting findings
But:
mathematically complex
requires strong statistical knowledge
USING STATISTICAL SOFTWARE
PACKAGES
Examples:
SPSS
SAS
Minitab
SYSTAT
Statistica
Advantages of Statistical Software
1. Wide Range of Statistics
Can perform:
advanced analyses
handle large datasets
manage missing data
2. User-Friendly
Results displayed in tables and graphs.
3. Assumption Testing
Software checks:
skewness
kurtosis
normality
4. Graphics
Creates:
charts
graphs
tables
IMPORTANT WARNING
Software cannot think for the researcher.
Researchers must:
choose correct statistical tests
interpret results properly