F.Y.B.
A Psychology Semester I: Introduction to Statistics, Mithibai College of Arts (Autonomous),
Mumbai
INTRODUCTION TO STATISTICS
F.Y.B.A Psychology – Semester I – Unit 1
References:
Howell, D (2013). Statistical Methods for Psychology (Eighth Edition). CA: Cengage
Learning.
Ciccarelli, S. K., & White, J. N. (2018). Psychology.5th edition. New Jersey: Pearson
Education.
1. Basics of Statistics
1.1 Population (N): is the entire collection of events (e.g. students’ scores, people’s incomes,
rat’s running speed, etc.) in which the researcher is interested. The population can be of any
size. They can range from a relatively small set of numbers, which can be collected easily, to
a large but finite set of numbers, which would be impractical to collect in their entirety. They
can be an infinite set of numbers (e.g. all possible cartoon drawings that students could
theoretically produce), which would be impossible to collect. The population is represented
with N (capital alphabet N).
1.2 Sample (n): Unfortunately for us in Psychology, we are interested in usually very large
set of numbers, which is impossible to collect. Hence, we draw only a sample of observations
from that population. We use this sample to infer something about the characteristics of that
population. Assuming that the sample is truly random, we cannot only estimate certain
characteristics of the population, but we can also have a very good idea of how accurate our
estimates are. Sample size is represented by small letter n.
1.3 Parameter: A measure that refers to an entire population is called a parameter. E.g.
average self-esteem score. Parameters are the real entities of the variable of interest.
1.4 Statistics: The measure that is collected from a sample is called statistic. Statistic is an
estimation/guess of the reality. So, we infer something about the characteristics of the
population (parameters) from what we know about the characteristics of the sample
(statistics).
1.5 Variable: A variable is a property of an object or event that can take on different values.
For example, hair colour is a variable because it is a property of an object (hair) and can take
on different values (brown, black, blue, grey, etc.)
1.6 Independent Variable: is the variable that is deliberately manipulated or controlled in an
experiment. It is usually identified as the potential cause of a behaviour/phenomenon.
Independent variables maybe either quantitative or qualitative and discrete or continuous.
1.7 Dependent Variable: is the variable that the researcher is interested to study. It is
influenced by the manipulation of the independent variable and is identified as the effect of
the independent variable manipulation. Dependent variables are generally (but not always)
quantitative and continuous.
Page 1 of 9
F.Y.B.A Psychology Semester I: Introduction to Statistics, Mithibai College of Arts (Autonomous),
Mumbai
1.8 Discrete variable: take on only a limited number of values (e.g. gender, high-school class,
number of days, etc.)
1.9 Continuous variable: takes on any value between the lowest and highest points on a scale
(e.g. self-esteem score, age, etc.)
There are also four scales of measurement: nominal, ordinal, interval, and ratio
2. Descriptive statistics
Descriptive statistics describe the data in an organized manner. It gives meaning to the raw
data by organizing it in a useful manner. The first step to statistical analyses is descriptive
statistics.
Descriptive statistics includes the following:
1. Frequency distribution and graphical representation
2. Measures of Central Tendency (Mean, Median and Mode)
3. Measures of Variability (Range, Standard Deviation, z Scores)
2.1.1Frequency Distribution
A frequency distribution is a table or graph that shows how often different numbers, or
scores, appear in a particular set of scores. The raw data is first distributed in order of the
magnitude and number of times each score is recorded.
For example, you collect the following data from a sample of 30 people about the number of
glasses of water they drink in a day.
2, 4, 5, 9, 3, 4, 5, 6, 7, 7, 8, 9, 3, 10, 4, 5, 6, 7, 8, 8, 5, 4, 6, 7, 6, 6, 5, 8, 7, 6
This is the raw information you have collected through a simple survey. Now, you distribute
the data in an organized manner. Draw the following frequency table. Just by looking at this
table, it is clear that typical people drink between four and eight glasses of water a day.
Number of glasses per day Tally Frequency
1 0
2 | 1
3 || 2
4 |||| 4
5 ||||| 5
6 |||||| 6
7 ||||| 5
8 |||| 4
9 || 2
10 | 1
30
Page 2 of 9
F.Y.B.A Psychology Semester I: Introduction to Statistics, Mithibai College of Arts (Autonomous),
Mumbai
2.1.2 Graphical Representation
Tables can be useful, especially when dealing with small sets of data. Sometimes a more
visual presentation gives a better “picture” of the patterns in a data set, and that is when
researchers use graphs to plot the data from a frequency distribution. One common graph is a
histogram, or a bar graph. Another type of graph used in frequency distributions is the
polygon, a line graph.
2.1.3 Normal Distribution Curve
Frequency polygons allow researchers to see the
shape of a set of data easily. For example, the number
of people drinking glasses of water in FigureA.2 is
easily seen to be centered about six glasses (central
tendency) but drops off below four glasses and above
eight glasses a day (variability). Our frequency
polygon has a high point, and the frequency decreases
on both sides.
A common frequency distribution of this type is called
the normal curve. It has a very specific shape and is
sometimes called the bell curve/ symmetrical curve. It
is also called the Gaussian curve. he normal curve is
used as a model for many things that are measured,
such as intelligence, height, or weight, but even those
measures only come close to a perfect distribution (provided large numbers of people are
measured). One of the reasons that the normal curve is so useful is that it has very specific
relationships to measures of central tendency and a measurement of variability, known as the
standard deviation.
2.1 4. Skewed Distribution
Distributions aren’t always normal in shape. Some
distributions are described as skewed. This occurs
when the distribution is not even on both sides of a
central score with the highest frequency (like in our
Page 3 of 9
F.Y.B.A Psychology Semester I: Introduction to Statistics, Mithibai College of Arts (Autonomous),
Mumbai
example). Instead, the scores are concentrated toward one side of the distribution. For
example, what if a study of people’s water-drinking habits in a different class revealed that
most people drank around seven to eight glasses of water daily, with no one drinking more
than eight? See Figure A-4.
These are Skewed distributions. The word ‘skew’
literally means ‘bias’. Skewed distributions can be
positively or negatively skewed, depending on where the
scores are concentrated. A concentration in the high end
would be called negatively skewed. A concentration in
the low end would be called positively skewed. The
direction of the extended tail determines whether it is
positively (tail to right) or negatively (tail to left) skewed.
2.2 Measures of Central Tendency
A frequency distribution is a good way to look at a set of numbers, but there’s still a lot to
look at—isn’t there some way to sum it all up? One way to sum up numerical data is to find
out what a “typical” score might be, or some central number around which all the others seem
to fall.
This kind of summation is called a measure of central tendency, or the number that best
represents the central part of a frequency distribution. There are three different measures of
central tendency: the mean, the median, and the mode.
2.2.1 Mean: The most commonly used measure of central tendency is the mean. It is the
arithmetic average of a distribution of numbers. Statistics mean is represented as X̅
X̅ = ⅀X / N
⅀ is a symbol called sigma. It is a Greek letter and it is also called the summation sign.
X represents a score. Rucha’s grades are represented by X.
⅀X means add up or sum all the X scores or ⅀X=86+92+37+90=355.
N means the number of scores.
For example, Rucha’s grades on the tests she has taken so far are 86, 92, 87, and 90.
We divide the sum of the scores (SX) by N to get the mean.
X̅ =⅀X/N = 355/4 = 88.75
The mean is a good way to find a central tendency if the set of scores clusters around the
mean with no extremely different scores that is either far higher or far lower than the mean.
2.2.2. Median: the mean doesn’t work as well when there are extreme scores, as you would
have if only two students out of an entire class had a perfect score of 100 and everyone else
scored in the 70s or lower. The median is not affected by such extreme scores.
Page 4 of 9
F.Y.B.A Psychology Semester I: Introduction to Statistics, Mithibai College of Arts (Autonomous),
Mumbai
A median is the score that falls in the middle of an ordered distribution of scores. Half of the
scores will fall above the median, and half of the scores will fall below it. If the distribution
contains an odd number of scores, it’s just the middle number, but if the number of scores is
even, it’s the average of the two middle scores. The median is also the 50th percentile.
Before calculating the median, it is important to arrange the raw data in an ascending order
and then find the central data that divides the set of scores into exact half.
E.g. You have collected the following set of IQ scores of a few people:
Arrange the scores in ascending order:
95, 98, 100, 100, 100, 102, 102, 139, 150, 160.
The mean of this data would be 1146 / 10 = 114.6
But the median would be the average between 100 and 102 i.e. 100+102/2 = 101.
2.2.3 Mode: Mode is the most frequently occurring score
in the frequency distribution. The mode is the previous set
data would be 100, because that number appears more
times in the distribution than any other. Three people have
that score.
2.2.4 Measures of Central Tendency and the shape of
the distribution: When the distribution is normal or close
to it, the mean, median, and mode are the same or
very similar.
If the distribution is skewed, then the mean is pulled in the direction of the tail of the
distribution. The mode is still the highest point, and the median is between the two.
2.3 Measures of Variability
Descriptive statistics can also determine how much the scores in a distribution differ, or vary,
from the central tendency of the data. These measures of variability are used to discover how
“spread out” the scores are from each other. The more the scores cluster around the central
scores, the smaller the measure of variability will be, and the more widely the scores differ
from the central scores, the larger this measurement will be.
There are two ways that variability is measured. The simpler method is by calculating the
range of the set of scores, or the difference between the highest score and the lowest score in
the set of scores. The range is somewhat limited as a measure of variability when there are
Page 5 of 9
F.Y.B.A Psychology Semester I: Introduction to Statistics, Mithibai College of Arts (Autonomous),
Mumbai
extreme scores in the distribution. The other measure of variability that is commonly used is
the one that is related to the normal curve, the standard deviation.
2.3.1 Standard Deviation: is the square root of the average squared difference, or deviation,
of the scores from the mean of the distribution.
SD = √⅀(X- X̅)2 / N
Steps for calculating the standard deviation include:
i. Scores are given/collected (X)
ii. Find the mean (X̅)
iii. Subtract each score (X) from its mean (X̅). This is the deviation of the score from its
mean (X- X̅). Scores higher than the mean would deviate in positive while scores
lower than the mean would deviate negatively. If the deviations are added up and
averaged, the score will be zero. Hence, we need to get rid of the negative deviations.
In mathematics, such problems are solved by squaring the numbers.
iv. Square the deviations (X- X̅) 2
v. Sum all the squared deviations ⅀(X- X̅) 2
vi. Find the S.D. = √⅀(X- X̅) 2/ N
Refer to the data of IQ scores collected to calculate the median. Use the same data here and
calculate the S.D. of the data.
X X̅ (X- X̅) (X- X̅)2
95 114.6 -19.60 384.16
98 114.6 -16.60 275.56
100 114.6 -14.60 213.16
100 114.6 -14.60 213.16
100 114.6 -14.60 213.16
102 114.6 -12.60 158.76
102 114.6 -12.60 158.76
139 114.6 24.40 595.36
150 114.6 35.40 1253.16
160 114.6 45.40 2061.16
⅀(X- X̅) = 5526.4
2
⅀X = 1146
S.D.= √5526.4/10 = √552.64 = 23.508
Page 6 of 9
F.Y.B.A Psychology Semester I: Introduction to Statistics, Mithibai College of Arts (Autonomous),
Mumbai
Standard Deviation and the Normal Curve:
Suppose the mean of a sample is 100 and the SD is 15. The normal distribution curve looks
like Figure A.8. It is a bell curve. With a true normal curve, researchers know exactly what
percentage of the population lies under the curve between each standard deviation from the
mean. For example, notice that in the percentages in Figure A.8, one standard deviation
above the mean has 34.13 percent of the population represented by the graph under that
section. These are the scores between 100 and 115. One standard deviation below the mean
(−1) has exactly the same percent, 34.13, under that section—the scores between 85 and 100.
This means that 68.26 percent of the population falls within one standard deviation from the
mean, or one average “spread” from the centre of the distribution. Although the “tails” of
this normal curve seem to touch the bottom of the graph, in theory they go on indefinitely,
never touching the base of the graph. In reality, though, any statistical measurement that
forms a normal curve will have 99.72 percent of the population it measures falling within
three standard deviations either above or below the mean.
2.3.2 Z score: Because this relationship between the standard deviation and the nor-mal
curve does not change, it is always possible to compare different test scores or sets of data
that come close to a normal curve distribution. This is done by computing a z score, which
indicates how many standard deviations you are away from the mean in terms of the number
of standard deviations that exist between the mean and that score. Z scores help us find the
relative position of any individual score in a distribution. It is done by locating how far away
from the mean is the individual score (X) in terms of SD units. It is calculated by subtracting
the mean from your score and dividing by the standard deviation.
Z = X - X̅ / S.D
In the above example, if you had an IQ of 115, your z score would be 1.0. If you had an IQ of
70, your z score would be −2.0. So on any exam, if you had a positive z score, you did
relatively well. A negative z score means you didn’t do as well.
Page 7 of 9
F.Y.B.A Psychology Semester I: Introduction to Statistics, Mithibai College of Arts (Autonomous),
Mumbai
3. Inferential statistics
After we have described our data in detail and have obtained basic information about the
data, we conduct inferential statistics to infer about the population from the sample data. Statistics
Since we infer about the population through the sample statistics, the concept of probability
must be understood.
Inferential Statistics are employed to determine what inferences or conclusions can
legitimately be drawn from a set of research findings. They are mathematical methods used to
determine how likely it is that a study’s outcome is due to chance and whether the outcome
can be legitimately be generalized to a larger population.
Depending on the data collected and objectives of research study, various inferential statistics
can be employed. If certain assumptions about the data are met, then parametric inferential
statistics can be employed. For example, Pearson’s Product Moment Correlations, the t-test(
to compare the means of two groups), F-test (also called Analysis of Variance or ANOVA,
employed to compare means of more than two groups), etc. However, if the assumptions
about the data are not met, then non-parametric inferential statistics are employed. For
example, Chi-square test, Kruskal-Wallis Test, Mann-Whitney U Test, etc.
Each inferential statistic helps determine how likely a particular finding has occurred as a
matter of chance/random variation.
If the findings suggest that the odds (likelihood/probability) of a particular finding occurring
are greater than mere chance, it can be concluded that the results are statistically significant.
We can infer with greater confidence that the manipulation of the independent variable (and
not simply chance factors) is the reason for the results.
This is done by finding the probability on the normal distribution curve and is expressed as
‘p’. No test can tell us for sure that the intervention/manipulation/IV has worked. It is
expressed as probability and not certainty. Some conventions have been developed to guide
researchers in their decisions about whether their results are statistically significant or not.
When the probability of obtaining a particular result (if random factors operate alone) is
LESS THAN 0.05 (i.e. 5 out of 100 cases), the result is considered statistically significant at
95% level. This means that 5 out of 100 cases would show result due to chance factors and
that there is 95% probability that the obtained result is due to experimental manipulation. It is
expressed as p<0.05.
When the probability of obtaining a particular result (if random factors operate alone) is
LESS THAN 0.01 (i.e. 1 out of 100 cases), the result is considered better statistically
significant at 99% level. This means that 1 out of 100 cases would show result due to chance
factor and that there is 99% probability that the obtained result is due to experimental
manipulation. Such a result is expressed as p<0.01.
In psychological research, usually these 2 levels of significance are employed in hypothesis
testing (0.05 level and 0.01 level).
Page 8 of 9
F.Y.B.A Psychology Semester I: Introduction to Statistics, Mithibai College of Arts (Autonomous),
Mumbai
Example: We collect the scores in Psychology from Division A, B and C. We find the means
of the three groups and run F-test to compare the statistical difference between the 3 groups.
Suppose we get the result as F=2.46, p<0.05. This can be interpreted as (a) There is a
statistical difference between the psychology scores of the three divisions (b) This statistical
difference is a true difference between the 3 divisions and is not due to some random chance
factors.
There are broadly two types of hypothesis employed:-
1. Null Hypothesis (HO) – means that there is no significant difference between the groups
compared. This means at M1 = M2. This is the hypothesis that is tested for rejection first by a
researcher.
2. Alternate Hypothesis (HA) / (H1) – means that there is a significant difference between the
groups compared. This means that M1≠ M2. It can be that M1>M2 or that M1<M2.
Page 9 of 9