M.A.
PSYCHOLOGY (Regular Mode)
AY 2025-2026 | Semester II | Course No. I (CORE)
STATISTICS IN PSYCHOLOGY (Course Code-201)
MODULE – I: INTRODUCTION TO STATISTICS IN PSYCHOLOGY
Detailed Study Notes
Source: Statistical Methods for Psychology – David C. Howell (7th Ed.)
MODULE I – TOPIC OVERVIEW
This module covers the foundational concepts of statistics as applied to psychological research.
The five major topics are:
• Scope of Statistics and Basic Concepts (Types of Data, Types of Variables, Population
& Sample)
• Measurement: Concept and Levels of Measurement
• Descriptive Statistics: Measures of Central Tendency and Dispersion
• Graphical Representation of Data
• Introduction to SPSS
TOPIC 1: Scope of Statistics and Basic Concepts
1.1 What is Statistics?
Statistics is the science of collecting, organising, analysing, and interpreting numerical data. In
psychology, it serves two broad purposes:
• Descriptive Statistics – summarising and describing a dataset (e.g., calculating means,
drawing graphs).
• Inferential Statistics – drawing conclusions about a population based on data from a
sample.
📝 Example: A researcher studies the effect of a stress-management programme on students' self-
esteem. She cannot test every student in the country, so she measures a sample and uses inferential
statistics to generalise the results.
1.2 Population and Sample
Population The entire collection of events or scores that the researcher is
interested in. It can be finite (e.g., all students in one school) or
theoretically infinite (e.g., all possible anxiety scores).
Sample A subset of the population that is actually measured. Because
populations are usually very large, researchers almost always work
with samples.
Random Sample A sample drawn so that every element in the population has an
equal chance of being selected. True random samples ensure
external validity.
Random Assignment Once participants are selected, they are randomly allocated to
treatment groups. This ensures internal validity — differences
between groups are due to the treatment, not pre-existing
differences.
Parameter A numerical value (e.g., mean, variance) that describes the entire
population. Denoted by Greek letters: μ (mean), σ² (variance).
Statistic A numerical value calculated from a sample, used to estimate the
corresponding population parameter. Denoted by Roman letters: X̄
(mean), s² (variance).
External Validity: Whether the sample accurately reflects the population so that findings can be
generalised.
Internal Validity: Whether differences in outcomes between groups are truly caused by the
treatment, not by how participants were assigned.
1.3 Types of Variables
A variable is any property of an object or event that can take on different values.
A. Independent vs Dependent Variables
Independent Variable (IV) The variable the researcher controls or manipulates (e.g., whether a
student received stress-management training). The 'cause'.
Dependent Variable (DV) The variable the researcher measures to see the effect (e.g., self-
esteem scores). The 'outcome'. Both start with 'd' — a helpful
memory cue.
Extraneous Variable Any variable other than the IV that could affect the DV. Must be
controlled or held constant.
Confounding Variable An extraneous variable that varies systematically with the IV,
making it impossible to isolate the IV's true effect.
Intervening Variable A variable that mediates the relationship between the IV and DV
(occurs in the causal pathway between them).
B. Continuous vs Discrete Variables
Continuous Variable Can take any value (including decimals) within a range. E.g., age,
reaction time, weight. In theory, any value between extremes is
possible.
Discrete Variable Can only take specific, separate values — usually whole numbers.
E.g., number of correct answers, gender, year in school.
1.4 Types of Data
Data can be classified along two dimensions: (a) what was measured and (b) how it was
measured.
A. Quantitative vs Qualitative (Categorical) Data
Quantitative / Numerical scores that indicate 'how much' of something. E.g., IQ
Measurement Data score = 115, reaction time = 640 ms. Each participant gets a unique
number.
Qualitative / Categorical / Data that classify participants into categories; results are counts
Frequency Data (frequencies). E.g., 34 females and 26 males; 15 'highly anxious',
33 'neutral'.
📝 The same underlying variable (e.g., Anxiety) can produce quantitative data if measured on a
continuous scale, or categorical data if people are sorted into groups (Low / Medium / High).
B. Continuous vs Discrete Data
• Continuous data arise from measuring a continuous variable (height, time).
• Discrete data arise from counting (number of children in a family, number of errors).
TOPIC 2: Measurement — Concept and Levels of
Measurement
2.1 What is Measurement?
Measurement is the process of assigning numbers to objects or events according to a set of
rules. The type of scale used determines what mathematical operations are permissible and
which statistical tests are appropriate.
S.S. Stevens (1951) classified measurement into four hierarchical levels, each with more
mathematical properties than the previous:
2.2 The Four Scales of Measurement
① Nominal Scale
The most basic level. Numbers (or labels) are assigned to categories purely as names — they
have no mathematical meaning whatsoever.
• Properties: Identity only. Category A is different from category B; that is all.
• Operations allowed: Count frequencies; find the mode.
• Psychology examples: Gender (1=Male, 2=Female), diagnosis (1=Depression,
2=Anxiety, 3=PTSD), jersey numbers in football.
📝 The numbers on football jerseys are a classic example: No. 7 is not 'greater' than No. 3 — the
numbers are just convenient labels.
② Ordinal Scale
Scores can be ranked from lowest to highest, but the intervals between ranks are not equal.
• Properties: Identity + Order.
• Operations allowed: Rank, find median, compute percentiles.
• Psychology examples: Military rank (Commander < Captain < Rear Admiral); finishing
position in a race (1st, 2nd, 3rd); Holmes-Rahe Stress Scale (a higher score = more
stress, but the gaps between scores are unequal).
📝 A person with stress score 20 has experienced more stress than someone scoring 10, but the
difference between 10 and 15 is not necessarily the same as between 15 and 20.
③ Interval Scale
Equal intervals between scale points are guaranteed. However, there is no true zero point —
zero is arbitrary.
• Properties: Identity + Order + Equal Intervals.
• Operations allowed: Addition, subtraction; compute mean and standard deviation.
• Psychology examples: Fahrenheit or Celsius temperature (the gap between 10°F and
20°F equals the gap between 80°F and 90°F); most psychological test scores (IQ,
personality inventories) are treated as interval.
• Limitation: Cannot form meaningful ratios. We cannot say 40°F is 'twice as warm' as
20°F.
📝 Converting Fahrenheit to Celsius changes the ratios, which shows the ratios were arbitrary in the
first place — a key indicator of an interval scale.
④ Ratio Scale
The highest level. Has all properties of interval scales plus an absolute (true) zero point,
meaning zero represents the complete absence of the attribute.
• Properties: Identity + Order + Equal Intervals + True Zero.
• Operations allowed: All arithmetic operations, including ratios.
• Psychology examples: Reaction time, number of correct responses, weight, length,
duration. '10 seconds is twice as long as 5 seconds' is a valid statement.
📝 Anxiety questionnaire scores are usually ordinal or interval at best — a score of 0 rarely means
'zero anxiety', so they typically lack a true zero.
2.3 Summary Comparison of Scales
Scale Properties Permissible Statistics Psychology Example
Nominal Identity Mode, frequency Gender, diagnosis
Ordinal + Order Median, rank Class rank, Likert-type
items
Interval + Equal intervals Mean, SD, correlation IQ scores, temperature
Ratio + True zero All statistics, ratios Reaction time, weight
2.4 Practical Importance of Scales
The measurement scale determines which statistical procedures are appropriate. However,
many researchers treat ordinal data (such as Likert scale responses) as interval when the
number of scale points is large enough and the distribution is approximately normal. The key
principle is that the scale reflects the underlying construct being measured, not just the numbers
assigned.
TOPIC 3: Descriptive Statistics
Descriptive statistics are used to summarise and describe the features of a dataset. They
reduce a large mass of numbers to a few interpretable values. Two fundamental aspects are
described: (a) central tendency — where the data cluster — and (b) dispersion — how spread
out the data are.
3.1 Measures of Central Tendency
Central tendency (also called measures of location) refers to a single value that best represents
the 'centre' of a distribution. The three main measures are the Mode, Median, and Mean.
A. Mode (Mo)
Definition The most frequently occurring score in a dataset. The value of X
with the highest frequency.
When to use Nominal data; to identify the most common category; when the
distribution is bimodal.
Advantages Simple to find; not affected by extreme scores; the only measure
valid for nominal data; the most probable value.
Disadvantages May not represent the whole dataset; can be unstable (changes
with grouping intervals); not usable in algebraic formulas.
Example In a dataset of test scores: {55, 70, 70, 85, 90, 90, 90, 95}, Mode =
90.
📝 A distribution with two modes is called bimodal; with one mode, unimodal. If two non-adjacent
values occur equally often, both are reported.
B. Median (Mdn)
Definition The middle score when data are arranged in ascending order. The
point at or below which 50% of scores fall. Equivalent to the 50th
percentile.
Formula (location) Median Location = (N + 1) / 2. For odd N: median is the middle
value. For even N: median is the average of the two middle values.
When to use Ordinal data; skewed distributions; when outliers are present.
Advantages Not influenced by extreme scores (resistant); applicable to ordinal
data; doesn't assume equal intervals.
Disadvantages Does not use all data values; less stable across samples than the
mean; harder to use in equations.
Example Data: {3, 5, 7, 8, 15} → N=5, Location=(5+1)/2=3 → Median = 7.
Data: {3, 5, 7, 11, 14, 15} → Location=3.5 → Median = (7+11)/2 =
9.
C. Mean (X̄ )
Definition The arithmetic average. Sum of all scores divided by the number of
scores.
Formula X̄ = ΣX / N where ΣX = sum of all scores, N = number of scores
When to use Interval or ratio data; symmetric, roughly normal distributions.
Advantages Uses all data points; algebraically manipulable; most efficient and
unbiased estimator of the population mean (μ); most stable across
samples.
Disadvantages Affected by extreme scores (outliers); may not be an actual value in
the dataset; requires at least interval-level data.
Example Scores: {3, 5, 12, 5} → X̄ = (3+5+12+5)/4 = 25/4 = 6.25
Choosing the Right Measure
Situation Best Measure Reason Example
Nominal data Mode Only valid measure for Most common political
categories party
Ordinal / Skewed Median Resistant to outliers & Household income
data skew
Interval / Ratio, Mean Uses all data; most Exam scores, reaction time
symmetric efficient
Bimodal distribution Mode (both) Two peaks need two Age distribution in a class
modes
📝 When a distribution is perfectly symmetric and unimodal, all three measures coincide. The more
skewed a distribution, the more the mean is pulled toward the tail.
3.2 Measures of Dispersion (Variability)
Dispersion measures describe how spread out or clustered the scores are around the central
tendency. Two datasets can have the same mean but very different dispersions.
A. Range
Definition The difference between the maximum and minimum scores.
Formula Range = Maximum score − Minimum score
Advantage Simple and quick to compute.
Disadvantage Entirely determined by extreme values; heavily affected by outliers;
ignores all other data.
Example Scores: {10, 20, 30, 40, 100} → Range = 100 − 10 = 90
B. Average Deviation (Mean Absolute Deviation — MAD)
Definition The average of the absolute deviations of each score from the
mean.
Formula MAD = Σ|X − X̄ | / N
Note Simply taking the average of (X − X̄ ) yields zero because positive
and negative deviations cancel. Taking absolute values avoids this
problem.
Limitation Rarely used in inferential statistics; the variance and standard
deviation are far more commonly used.
C. Quartile Deviation (Semi-Interquartile Range)
Quartile (Q) Values that divide the sorted dataset into four equal parts: Q1 (25th
percentile), Q2 (median / 50th), Q3 (75th percentile).
Interquartile Range (IQR) IQR = Q3 − Q1. The range of the middle 50% of scores.
Advantage Not influenced by extreme values; useful for skewed data and
boxplots.
Disadvantage Discards the top and bottom 25% of data; less algebraically
tractable.
Example Attractiveness ratings: Q1=2.0, Q3=3.0 → IQR = 1.0
D. Variance (s²)
Definition The average of the squared deviations from the mean. Squaring
removes negative signs and amplifies larger deviations.
Sample Variance s² = Σ(X − X̄ )² / (N − 1) [divide by N−1 for an unbiased estimate of
Formula the population variance σ²]
Population Variance σ² = Σ(X − μ)² / N
Why N−1? Because X̄ is estimated from the data (1 degree of freedom is used
up), so only N−1 values are truly free to vary. Using N−1 makes s²
an unbiased estimator of σ².
Limitation Expressed in squared units (e.g., 'squared seconds'), which lacks
intuitive meaning.
Extreme scores Because deviations are squared, outliers have a disproportionately
large effect on the variance.
E. Standard Deviation (s or SD)
Definition The positive square root of the variance. Expressed in the same
units as the original data, making it interpretable.
Sample SD Formula s = √[Σ(X − X̄ )² / (N − 1)]
Interpretation Roughly, the 'average' amount by which individual scores deviate
from the mean. For a normal distribution, ≈68% of scores fall within
±1 SD of the mean; ≈95% within ±2 SD.
Population symbol σ (sigma) for population; s for sample.
Advantage Most widely used measure of spread; uses all data; algebraically
tractable.
Example If mean exam score = 70 and s = 10, approximately 68% of
students scored between 60 and 80.
Summary Table: Measures of Dispersion
Measure Formula Uses All Data? Sensitive to Outliers?
Range Max − Min No (2 values only) Very high
MAD Σ|X−X̄ |/N Yes Moderate
IQR Q3 − Q1 No (middle 50%) Low (resistant)
Variance (s²) Σ(X−X̄ )²/(N−1) Yes High (squares)
Std. Dev. (s) √[Σ(X−X̄ )²/(N−1)] Yes High
TOPIC 4: Graphical Representation of Data
Graphs and plots allow researchers to see the shape, centre, spread, and outliers of a dataset
at a glance. Different graphs suit different types of data and research questions.
4.1 Frequency Distribution
Definition A table (or graph) that lists each possible score and the number of
times it occurred. The most basic form of data organisation.
Simple Frequency Lists each raw score with its frequency. Useful when few distinct
Distribution values exist.
Grouped Frequency Data are grouped into class intervals (e.g., 35–39, 40–44). About 10
Distribution intervals are usually recommended. Useful when data range is
large.
Cumulative Frequency Running total of frequencies from lowest to highest interval; shows
how many cases fall at or below each interval.
Real Limits The true boundaries of an interval lie halfway between adjacent
intervals. E.g., 35–39 has real limits 34.5–39.5.
Midpoint The average of the upper and lower limits of an interval; used when
plotting grouped data.
4.2 Histogram
Definition A bar chart for continuous data where bars touch each other (unlike
discrete bar charts). The X-axis represents scores or intervals; Y-
axis represents frequency.
Construction Draw adjacent bars for each interval; bar height = frequency of that
interval; no gaps between bars (emphasises continuity).
Advantage Shows the shape of the distribution: symmetric, skewed, bimodal,
etc.
When to use Quantitative continuous data (exam scores, heights, reaction
times).
Note on intervals Choosing different interval widths can make the same dataset look
very different. Use natural break points and aim for ~10 intervals.
Describing Distribution Shape
Symmetric Mirror image on both sides of the centre. Mean = Median = Mode
(for unimodal symmetric distributions).
Positively Skewed Tail extends to the right (high values). Mean > Median > Mode.
E.g., income distribution.
Negatively Skewed Tail extends to the left (low values). Mean < Median < Mode. E.g.,
easy exam scores.
Unimodal One clear peak.
Bimodal Two prominent peaks.
Kurtosis How peaked vs. flat a distribution is. Mesokurtic (normal),
Leptokurtic (very peaked, heavy tails), Platykurtic (flat, thin tails).
Outlier A score that is extremely distant from the rest of the distribution.
May indicate a data-entry error or a rare genuine observation.
4.3 Bar Graph
Definition A graph with separate, non-touching bars, one per category. The
height of each bar represents the frequency or percentage of that
category.
Difference from Bars do NOT touch (emphasising discrete / categorical nature of
Histogram data). Used for nominal or ordinal data.
When to use Categorical / nominal data (e.g., number of males vs. females;
frequency of each diagnosis).
Example A bar graph showing the number of participants in each year of
study (1st year, 2nd year, 3rd year).
4.4 Line Graph (Frequency Polygon)
Definition Points are plotted at midpoints of each interval and connected by
straight lines. Anchored to the baseline at both ends.
Advantage Better for showing trends over time; allows overlaying two or more
distributions for comparison.
When to use Continuous data; when comparing two distributions on the same
axes.
Frequency Polygon A line graph of a frequency distribution; the area under the polygon
equals the total frequency.
4.5 Pie Chart
Definition A circle divided into segments, where each segment's size is
proportional to the percentage of the whole that category
represents.
When to use To show the relative proportions of categories; especially effective
when there are few categories and the emphasis is on part-to-whole
relationships.
Limitation Difficult to read when there are many small segments; not suitable
for showing changes over time.
Example A pie chart showing the percentage of psychology students in
different specialisation tracks.
4.6 Box-and-Whisker Plot (Boxplot)
Definition A compact graph showing the five-number summary: Minimum, Q1,
Median (Q2), Q3, and Maximum, plus outliers.
Construction Draw a box from Q1 to Q3; draw a vertical line at the median inside
the box; draw 'whiskers' from Q1 down to the lower adjacent value
and from Q3 up to the upper adjacent value; plot outliers
individually as separate points.
Inner Fence Q1 − 1.5×IQR (lower) and Q3 + 1.5×IQR (upper). Values beyond
the fences are outliers.
Adjacent Values The most extreme actual data values that still fall within the inner
fences.
Advantages Displays centre, spread, skewness, and outliers simultaneously;
very effective for comparing two or more groups.
Identifying Skewness If the median is closer to Q1, the data are positively skewed; if
closer to Q3, negatively skewed. A longer upper whisker also
signals positive skew.
Choosing the Right Graph
Graph Type Data Type Primary Purpose Example Use
Bar Graph Nominal / Ordinal Compare categories Frequency of diagnoses
Histogram Continuous Show distribution shape Distribution of exam scores
(Interval/Ratio)
Frequency Polygon Continuous Show shape; compare Overlay two score
groups distributions
Pie Chart Nominal Part-to-whole relationships % students in each major
(proportions)
Line Graph Continuous over Show trends Scores over 5 sessions
time
Box Plot Interval / Ratio Centre, spread, outliers Compare reaction times by
group
TOPIC 5: Introduction to SPSS
SPSS (Statistical Package for the Social Sciences) is one of the most widely used statistical
software packages in psychological research. It provides a user-friendly, menu-driven interface
for managing data, running analyses, and creating graphs.
5.1 Data Editor
When you open SPSS, the first window you see is the Data Editor. It has two views:
A. Data View
Appearance Looks like a spreadsheet. Each row = one case (participant). Each
column = one variable.
Entering data Click on a cell and type the value. Press Enter or Tab to move to
the next cell.
Missing values Empty cells are treated as missing. SPSS represents system-
missing values as a dot (.). User-defined missing values can be
specified in Variable View.
B. Variable View
Purpose Define the properties of each variable before or after entering data.
Name A short label for the variable (no spaces; max 64 characters). E.g.,
'AnxScore'.
Type The data type: Numeric (default), String (text), Date, etc.
Width & Decimals Number of digits and decimal places to display.
Label A longer, descriptive name shown in output. E.g., 'Anxiety
Questionnaire Score'.
Values Assign numeric codes to category labels. E.g., 1 = 'Male', 2 =
'Female'.
Missing Specify numeric codes that represent missing data (e.g., 99 means
'no response').
Measure Identify the scale of measurement: Nominal, Ordinal, or Scale
(Interval/Ratio).
5.2 SPSS Viewer (Output Window)
All results — tables, charts, and notes — appear in the SPSS Viewer (Output Window) after an
analysis is run. Key features:
• Navigation tree on the left: Click any item to jump to that output.
• Results can be double-clicked to edit, right-clicked to copy, or exported to Word or PDF.
• Pivot Tables: SPSS output tables can be pivoted (rows/columns swapped) and
formatted.
• Log section: Shows the syntax that was executed, useful for replication.
5.3 Obtaining Descriptive Statistics in SPSS
Menu path: Analyze → Descriptive Statistics → Frequencies (or Descriptives)
• Frequencies: Provides frequency tables, measures of central tendency and dispersion,
and graphs (histograms, bar charts, pie charts).
• Descriptives: Provides mean, SD, range, min, max, and z-scores for each variable.
• Explore: Provides a complete suite including stem-and-leaf plots and boxplots — very
useful for checking distributional assumptions.
• Compare Means → Means: Computes descriptive statistics broken down by group
levels.
Typical output from Analyze → Descriptive Statistics → Frequencies includes:
◦ N: Number of valid cases and number missing
◦ Mean: Arithmetic average
◦ Median: Middle value (50th percentile)
◦ Mode: Most frequent value
◦ Std. Deviation: Sample standard deviation
◦ Variance: Sample variance (SD²)
◦ Range: Max − Min
◦ Minimum / Maximum: Smallest and largest values
5.4 Creating Graphs in SPSS
Legacy Dialogs Graphs → Legacy Dialogs → Histogram / Bar / Line / Pie / Boxplot.
Provides quick, basic graphs.
Chart Builder Graphs → Chart Builder. Drag-and-drop interface for creating
customised charts.
Interactive / Available from SPSS 25+ for more modern chart styles.
Visualizations
Editing Charts Double-click any chart in the Output Viewer to open the Chart Editor
for formatting colours, fonts, and axis labels.
5.5 Saving and Retrieving Files
Data file (.sav) Save with File → Save As → SPSS Statistics (*.sav). Contains the
dataset and variable definitions.
Output file (.spv) Save with File → Save → SPSS Viewer (*.spv). Contains all your
results, tables, and graphs.
Syntax file (.sps) SPSS commands written as text. Save with File → New → Syntax,
then File → Save. Allows reproducible analyses.
Opening files File → Open → Data (for .sav) or Output (for .spv). SPSS can also
import Excel, CSV, and other formats.
Exporting output File → Export → output can be saved as PDF, Word (.docx), Excel
(.xlsx), HTML, or PowerPoint.
📝 Best practice: Always save both the .sav data file and the .sps syntax file. The syntax file is a
record of every step of your analysis and allows you to reproduce or modify results later.
QUICK REVISION — Key Formulas & Definitions
Concept Formula / Key Point
Mean (X̄ ) X̄ = ΣX / N
Median Location (N + 1) / 2
Sample Variance (s²) s² = Σ(X − X̄ )² / (N − 1)
Population Variance (σ²) σ² = Σ(X − μ)² / N
Standard Deviation (s) s = √[Σ(X − X̄ )² / (N − 1)]
Range Max score − Min score
IQR Q3 − Q1
Lower Inner Fence Q1 − 1.5 × IQR
(boxplot)
Upper Inner Fence Q3 + 1.5 × IQR
(boxplot)
Coefficient of Variation CV = (s / X̄ ) × 100
(CV)
Degrees of Freedom (df) N − 1 (for single-sample variance)
— End of Module I Notes —
Based on: Statistical Methods for Psychology, 7th Edition — David C. Howell | Osmania University M.A. Psychology
Curriculum