Quantitative Methods in Psychology PSY103
Quantitative Methods in Psychology PSY103
PSY103
All rights reserved. No part of this publication may be reproduced, stored in a retrieval
system, or transmitted in any form or by any means, electronic, mechanical,
photocopying, recording or otherwise, without the prior permission of the copyright
owner.
ISBN: 978-021-278-7
???????
PSY103 Quantitative Methods in Psychology
Contents
Course overview 3
Welcome to Quantitative Methods in Psychology PSY103 .................................. 3
Quantitative Methods in Psychology PSY103—is this course for you? ............... 3
Timeframe .............................................................................................................. 4
Study skills ............................................................................................................. 4
Need help? ..................................................... ????Error! Bookmark not defined.
Assessments .................................................. ????Error! Bookmark not defined.
Study Session 1 8
The Meaning of Statistics....................................................................................... 8
Introduction ................................................................................................... 8
1.1 The Definition of Statistics...................................................................... 8
1.1.1 Layman‘s Perspective .................................................................. 9
1.1.2 Research Perspective ................................................................... 9
1.1.3 Grammatical Perspective ........................................................... 10
1.2 Types of Statistics ................................................................................. 10
1.2.1 Descriptive Statistics ................................................................. 11
1.2.2 Inferential Statistics ................................................................... 11
Study Session Summary ....................................................................................... 13
Assessment ........................................................................................................... 13
Study Session 2 14
Why we Study Statistics....................................................................................... 17
Introduction ................................................................................................. 17
Rationale for the Inclusion of Statistics into Psychology ........................... 17
Study Session Summary ....................................................................................... 19
Assessment ........................................................................................................... 19
Study Session 3 19
Common Terms used in Statistics ........................................................................ 20
Introduction ................................................................................................. 20
3.1 Variable ................................................................................................. 20
Contents
Study Session 4 29
Sampling............................................................................................................... 29
Introduction ................................................................................................. 29
4.1 What is Sampling? ................................................................................. 29
4.1.1 Important Concepts in Sampling ............................................... 29
4.2 Sampling Techniques ............................................................................ 30
4.2.1 Probability Sampling Techniques .............................................. 30
A. Simple Random Sampling ..................................................... 30
B. Stratified Random Sampling .................................................. 32
C. Cluster Sampling .................................................................... 33
D. Systematic Sampling .............................................................. 33
4.2.2 Non-probability Sampling ......................................................... 34
A. Accidental Sampling .............................................................. 34
B. Quota Sampling ...................................................................... 34
C. Purposive Sampling................................................................ 34
Study Session Summary ....................................................................................... 37
Assessment ........................................................................................................... 37
Study Session 5 39
Scales / Levels of Measurement ........................................................................... 39
Introduction ................................................................................................. 39
5.1 Nominal Scale ....................................................................................... 39
5.2 The Ordinal Scale .................................................................................. 40
5.3 Interval Scale ......................................................................................... 41
5.4 Ratio Scale ............................................................................................. 42
Study Session Summary ....................................................................................... 43
Assessment ........................................................................................................... 43
Study Session 6 45
Frequency Distribution......................................................................................... 45
Introduction ................................................................................................. 45
6.1 Regular Frequency Distribution ............................................................ 45
PSY103 Quantitative Methods in Psychology
Study Session 7 56
Graphic Representation of Frequency Distribution ............................................. 56
Introduction ................................................................................................. 56
7.1 The Frequency Polygon......................................................................... 56
8.2 The Histogram ....................................................................................... 57
7.3 The Bar Chart ........................................................................................ 59
Study Session Summary ....................................................................................... 61
Assessment ........................................................................................................... 61
Study Session 8 63
Diagrammatic Representation of Frequency Distribution ................................... 63
Introduction ................................................................................................. 63
8.1 The Pie Chart ......................................................................................... 63
8.2 The Pictogram ....................................................................................... 66
8.3 The Stem-and-Leaf Display .................................................................. 67
Study Session Summary ....................................................................................... 69
Assessment ........................................................................................................... 69
Study Session 9 71
Cumulative Distributions ..................................................................................... 71
Introduction ................................................................................................. 71
9.1 Constructing Cumulative Distribution .................................................. 71
9.1.1 Constructing a ―More Than‖ Cumulative Distribution ............. 71
9.1.2 Constructing a ―Less Than‖ Cumulative Frequency Distribution
............................................................................................................ 72
9.2 Graphical Presentations of Cumulative Distribution ............................ 73
9.2.1 Graphical Presentation of a ―More Than‖ Cumulative
Distribution. ........................................................................................ 73
9.2.2 Graphical Presentation of ―Less Than‖ Cumulative Distribution.
............................................................................................................ 74
Contents
Study Session 10 77
Measures of Central Tendency:Mean .................................................................. 77
Introduction ................................................................................................. 77
10.1 The Arithmetic Mean .......................................................................... 78
10.1.1 Ungrouped Data ....................................................................... 78
10.1.2 Regular Frequency Distribution .............................................. 79
10.1.3 Grouped Frequency Distribution with Class Intervals ............ 80
..................................................................... Error! Bookmark not defined.
Study Session Summary ....................................................................................... 81
Assessment ........................................................................................................... 82
Study Session 11 82
Measures of Central Tendency: Median .............................................................. 83
Introduction ................................................................................................. 83
11.1 The Median of Ungrouped Data .......................................................... 83
11.2 The Median of Grouped Data .............................................................. 84
11.2.1 The Median of Grouped Data with Regular Frequency
Distribution ......................................................................................... 84
11.2.2 The Median of Grouped Data with Grouped Frequency
Distribution with Class Intervals ........................................................ 85
11.3 Determining the Median using the Cumulative Frequency (Or OGIVE
Curve) .......................................................................................................... 88
Study Session Summary ....................................................................................... 89
Assessment ........................................................................................................... 89
Study Session 12 90
Measures of Central Tendency: The Mode .......................................................... 90
Introduction ................................................................................................. 90
12.1 Mode of Grouped Data with Intervals ................................................. 90
12.1.1 The Crude Mode ...................................................................... 90
12.1.2 The Interpolated Mode ............................................................ 91
12.2 Determining the Mode using the Histogram ....................................... 93
12.3 The Characteristics and Use of Measures of Central Tendency ......... 93
Study Session Summary ....................................................................................... 95
Assessment ........................................................................................................... 95
Study Session 13 96
Measures of Variability ........................................................................................ 96
Introduction ................................................................................................. 96
PSY103 Quantitative Methods in Psychology
References 115
document.
Quantitative Methods in Psychology PSY103 has been
produced by University of Ibadan Distance Learning Centre.
All Psychology Error! No text of specified style in document.s
produced by University of Ibadan Distance Learning Centre
are structured in the same way, as outlined below.
Your comments
After completing this course, Quantitative Methods in
Psychology, we would appreciate it if you would take a few
moments to give us your feedback on any aspect of this
course. Your feedback might include comments on:
Course content and structure.
Course reading materials and resources.
Course assessments.
Course assignments.
Course duration.
Course support (assigned tutors, technical help, etc).
Your general experience with the course provision as a
distance learning student.
You might forward your comments to
coursereview@[Link] or post same on comment pad of
your course website at UI Mobile Class. Your constructive
feedback will help us to improve and enhance this course.
2
PSY103 Quantitative Methods in Psychology
Course overview
Timeframe
How long?
Study skills
As an adult learner your approach to learning will be different
to that from your school days: you will choose what you want
to study, you will have professional and/or personal
motivation for doing so and you will most likely be fitting
your study activities around other professional or domestic
responsibilities.
Essentially you will be taking control of your learning
environment. As a consequence, you will need to consider
performance issues related to time management, goal setting,
stress management, etc. Perhaps you will also need to
reacquaint yourself in areas such as essay planning, coping
with exams and using the web as a learning resource.
Your most significant considerations will be time and space
i.e. the time you dedicate to your learning and the
environment in which you engage in that learning.
We recommend that you take time now—before starting your
self-study—to familiarize yourself with these issues. There are
a number of excellent web links and resources on the ―Self-
4
PSY103 Quantitative Methods in Psychology
Need help?
Help You may contact any of the following units for information,
learning resources and library services.
Academic Support
Activities
Assessments
Bibliography
6
PSY103 Quantitative Methods in Psychology
Study Session 1
Introduction
Statistics is perceived and interpreted differently by different
professionals. This study session will therefore expose you to
alternative definitions and explanations of the concept of
statistics.
When you have studied this session, you should be able to:
explain the concept of statistics, in grammatical terms,
and as used by scientific researchers. (SAQ 1.1)
Learning
Outcomes explain the two basic types of statistics. (SAQ1.2)
Yes
No
2. Research perspective
3. Grammar perspective
1.1.1 Layman’s Perspective
As used in everyday language, statistics implies a collection of
numerical data. Whenever figures or numerical data are used to
represent observations, it is called statistics. This may be
expressed in terms of the number of students in a school, the
number of employees in a company, the number of children given
birth to in a year, the number of accidents recorded in a year and
so forth.
o ITQ In a lay-person‘s understanding, statistics
means:
A. Addition, division, and subtraction of data.
B. A collection of numerical data.
Feedback on ITQs answers
The correct answer was B
If you have chosen A (addition, division and
subtraction of data), then you have defined based on
mathematical experience in primary school. In fact
the question asks you to define statistics as a
collection of numerical data.
1.1.2 Research Perspective
The most common inferential statistical tests include the student's t-test
and the Analysis of Variance (ANOVA) or F-test; these statistics are
used to determine whether the average differences across comparison
groups are due to the effects of the independent variable under study.
Another commonly used inferential statistic is the correlation
coefficient, which describes the strength of the relationship between
Study Session 1 The Meaning of Statistics
You also learnt about types of statistics which are: descriptive and
inferential statistics. The relationship between these types of
statistics was also highlighted.
Assessment
Now that you have completed this study session, you can
assess how well you have achieved its Learning Outcomes by
answering these questions. Write your answers in your Study
Assessment Diary and discuss them with your Tutor at the next Study
Support Meeting. You can check your answers with the Notes
on the Self-Assessment Questions at the end of this Manual.
Study Session 2
History of Statistics
Introduction:
This Study Session is concerned with the history of Statistics. Most
times when statististics is taught, especially outside the scope of its
domain, not much is said about its history. A body of knowledge is
better appreciated when its history is known. It is on this basis that this
Study Session is carved out for knowledge
Learning Outcomes:
At the end of this Study Session, the learner would be able to achieve:
(a) Additional knowledge of the definition of statistics from
historilcal background
(b) Gain knowledge on the contribution of some individuals
in the origin of statistics
(c) Appreciate areas of historical background in the use of
statistics
During the 20th century, the field led to the creation of precise
instruments for public health concerns, especially in the area of
epidemiology, biostatistics, and the like, and economic and social
purposes like unemployment rate, econometry, etc. All these
necessitated substantial advances in statistical practices. After World
War I, the necessity for Western welfare states to develop a specific
knowledge of their "population" promoted statistical concerns.
Philosophers like Michel Foucault argued that these developments
Assessment: The Meaning of Statistics
Today the use of statistics has broadened far beyond its origins as a
service to a state or government. Individuals and organizations use
statistics to understand data and make informed decisions throughout
the natural and social sciences, medicine, business, and other areas.
Statistics is generally regarded not as a subfield of mathematics but as a
distinct, allied, field. Many universities maintain separate mathematics
and statistics departments. Statistics is also taught in departments as
diverse as psychology, education, and public health.
Assessment:
a) Apart from the three definitional perspectives of
statistics, provide an historical definitional perspective
b) Mention the names of some founding fathers of statistics
c) Mention areas of human endeavours where statistics
could be applied.
PSY103 Quantitative Methods in Psychology
Study Session 3
Introduction
This Study Session aims to provide explanations on why students
of psychology and other related disciplines require the knowledge
of statistics.
When you have studied this session, you should be able to:
justify the inclusion of statistics in the study of
Learning psychology. (SAQ2.1)
Outcomes
Assessment
Now that you have completed this study session, you can
assess how well you have achieved its Learning Outcomes by
answering these questions. Write your answers in your Study
Assessment Diary and discuss them with your Tutor at the next Study
Support Meeting. You can check your answers with the Notes
on the Self-Assessment Questions at the end of this Manual
SAQ 2.1 (tests Learning Outcome 3.1)
Identify at least four reasons why it is necessary for you to
study statistics?
Study Session 4
Introduction
The teaching of statistics involves the understanding of concepts
that are used in scientific research. This study session will
therefore focus on scientific concepts that are used in statistics.
When you have studied this session, you should be able to:
explain the meanings of concepts such as variables,
Learning data, population, parameter, and sample. (SAQ4.1)
Outcomes
4.1 Variable
Variables are any measurable feature, which has the potential of
taking different values when it is quantified. It includes such
human features like weight, sex, intelligence quotient (I.Q), etc.
These human characteristics can assume different values when
different individuals are evaluated, or when the same individuals
are evaluated on different occasions.
4.1.1 Classification of Variables
Variables in research are classified into three:
Independent Variables (IV)
Independent variables are those that are easily manipulated by
researchers so as to determine their effects on another set of
variables called the dependent variables. In fact, the independent
variables are the variables which the researcher determines its
variations to examine how the outcome variable is affected. In
behavioral research, the independent variable usually consists
of the two (or more) treatment conditions to which participants
are exposed. The independent variable consists of the
antecedent conditions that were manipulated prior to observing
the dependent variable.
PSY103 Quantitative Methods in Psychology
4.2 Data
Data are information in Collections of numbers or measurements used to represent or
raw or unorganized form
(such as numbers or
quantify observations during research before transformation to
symbols).. other statistical forms is referred to as raw data. For instance,
students‘ scores in an examination can be used to represent
academic performance, and could be taken as data for the
research on effect of noise on academic performance.
4.3 Population
This is also referred to as universe. It is ideally defined as a large
group of people, animals, objects, responses, or measurements
that are alike in at least one respect. Thus, we could have a
population of Yoruba students in the University of Ibadan, a
population of all ‗white rats‘ of a given genetic strain, the
population of undergraduate students in the University of Nigeria,
and the population of all federal government workers in Ibadan.
4.3.1 Parameter
This refers to the describable features of a population. They are
the features or characteristics that a collection of people or objects
possess that allow them to be called a population. For example, a
population of undergraduate students has the feature of grouping
all undergraduates, and thereby excluding the postgraduate
students. We may have the male population in Ibadan also,
comprising all male persons in Ibadan.
4.3.2 Sample
Representative Sample: A This is a subset of a population, usually drawn from the
subset of a statistical
population that accurately
population randomly or otherwise, for the purpose of
reflects the members of the generalization to the population. In the behavioural sciences
entire population. and education, it is not often easy to study the entire
population of interest. Consequently, only a part of the
population is taken for study and this part is representative of
the entire members in the population under study. The part
selected for a study is the sample. When the sample is selected
in scientific ways, all the characteristics of the population are
believed to be reflected in the chosen sample. Such sample is
then referred to as a representative sample of the population.
In a class of 60 students, 10 of them could be selected as a
sample if the selection is done randomly. We will discuss the
concept of sampling in the next Study Session.
Once you have collected and recorded data in a table like you
have in questions 3 and 4 of the activity above, you have a
frequency table of your data. If you look at the frequency
table, you can start to answer descriptive and interpretive
Activity 3.2 questions about your data such as the questions in the
Allow 5 minutes following activity.
1. Which colour was the most common in your collection of
books?
2. Which colour was the least common in your collection of
books?
3. Do you think that another collection of books would have
the same numbers of different colours as you have found?
Assessment
Now that you have completed this study session, you can
assess how well you have achieved its Learning Outcomes by
answering these questions. Write your answers in your Study
Assessment Diary and discuss them with your Tutor at the next Study
Support Meeting. You can check your answers with the Notes
on the Self-Assessment Questions at the end of this Manual
SAQ 4.1 (tests Learning Outcome 4.1)
Study Session 4 Common Terms used in Statistics
Carefully study the scenario in the short passage below, and fill in
the blank spaces.
Mr. Biodun, a researcher studying depression gave a new
treatment to a sample of 100 people with depression out of
a____A____, a group of all things that share a set of
characteristics. In this case, the ―things‖ are people, and the
characteristic they all share is depression.
Mr. Biodun wants to know what the mean depression score for the
population would be if all people with depression were treated
with the new depression treatment. In other words, he wants to
know the____B___, the value that would be obtained if the entire
population were actually studied. Of course, Mr. Biodun don‘t
have the resources to study every person with depression in the
world, so he must instead study a____C___, a subset of the
population that is intended to represent the population. In most
cases, the best way to get a sample that accurately represents the
population is by taking a ___D___ from the population, when
each individual in the population has the same chance of being
selected for the sample.
So, he then uses the sample statistic value as an estimate of the
population parameter value. When researchers use a sample
statistic to infer the value of a population parameter it is called
___E___.
PSY103 Quantitative Methods in Psychology
Study Session 5
Sampling
Introduction
This Study Session will expose you to the meaning of sampling,
and the various techniques of conducting sampling, specifically in
social science research setting .
When you have studied this session, you should be able to:
use sampling technique. (SAQ 4.1)
Learning differentiate between probability and non-probability
Outcomes sampling techniques. (SAQ 4.1 and 4.2)
control group. The toss of a coin has been a method used to determine
random outcomes for centuries. It is still used in some research studies
as a method of randomization, although it has largely been discredited
as a valid randomization method. In football matches, it is used to
randomly determine what team starts the kick.
because they fit a particular profile. The researcher sets some criteria
based on the needs of the research, and selects from the population
those who fit into the set criteria. A researcher might, for example,
want to select his sample for a research; entrepreneurs who have a
minimum size of 10 employees. While the findings from purposive
sampling do not always have to be statistically representative of the
greater population of interest, they are qualitatively generalizable. The
more prior information researchers have about their particular
population of interest, the better the sample that they are going to
select.
D. Convenience Sampling
Convenience sampling is a non-probability sampling method where
participants are selected for inclusion in the sample because they are
the easiest for the researcher to access. This can be due to geographical
proximity, availability at a given time, or willingness to participate in
the research. Convenience sampling involves using respondents who
would be ―convenient‖ for the researcher to access. The researcher
leverage individuals that can be identified and approached with as little
effort as possible. Although there is no pattern whatsoever in acquiring
these respondents, but they may be recruited by merely asking people
who are present in the street, in a public building, or in a workplace, for
example, to participate in the research.
Assessment
Now that you have completed this study session, you can
assess how well you have achieved its Learning Outcomes by
Assessment answering these questions. Write your answers in your Study
Diary and discuss them with your Tutor at the next Study
Support Meeting. You can check your answers with the Notes
on the Self-Assessment Questions at the end of this Manual
SAQ 5.1 (tests Learning Outcome 5.1)
A researcher desires to select 200 participants for a research to
examine the relationship between sex and academic performance
at the University of Ibadan. She however discovers that the ratio
of males : females in the institution‘s sample frame is 3:2. What
number of males : females should be selected from this population
for the study?
A. 150:50
B. 80:120
C. 120 : 80
D. 50:150
SAQ 5.2 (tests Learning Outcome 5.1 and 5.2)
Assuming that a population size is known to be 600 and the above
researcher wishes to systematically sample 20 participants from
that population. Determine the nth term of the random selection
for the sample.
A. 20
B. 30
Study Session 5 Sampling
C. 600
D. 400
Find the first and last individuals that would be included in
the sample in SAQ 5.2. Post your answer on Study Session
4 Assignment Page at the UI Mobile Class.
Assignment
PSY103 Quantitative Methods in Psychology
Study Session 6
Introduction
Measurement scale is very pivotal in determining what statistical
analysis would be appropriate during an analysis. This Study
Session is therefore designed to acquaint you on various forms in
which data can be expressed in research.
When you have studied this session, you should be able to:
define , use correctly and differentiate all of the key
words printed in bold. (SAQ 6.1)
Learning Nominal scale
Outcomes
Ordinal scale
Interval scale
Ratio scale
To put data of a variable into statistical use, they are recorded in
some systematic ways. The procedure of describing the variables
is called scaling, and the descriptions which emerge are called
scales. The scales on which variables are measured have different
properties.
Scales of measurement are also referred to as levels of
measurement. This is because the numbers used to represent some
variables have more meaning than those used to represent other
variables.
Four scales are conventionally used in converting observations
into numerical data. These are nominal, ordinal, interval and ratio
scales.
between the 2nd and 3rd student. Although ordinal scales permit
the use of numbers with quantitative meaning, the numerals
employed are non-quantitative in that arithmetic computations of
addition, subtraction, multiplication and division are not
computable, just like the nominal scales. Variables expressed in
ordinal terms include job status, class levels at school, level of
education, etc.
o ITQ Religious grouping is what type of
measurement?.
A. Nominal scale.
B. Ordinal scale.
Feedback on ITQs answers
The correct answer is A, because religious grouping
is just a descriptive categorization.
If you chose B, you would be wrong because ordinal
scales are not mere categories, they also show
relative size.
Assessment
Now that you have completed this study session, you can
assess how well you have achieved its Learning Outcomes by
answering these questions.
Assessment
Pets: 5 dogs, 12
cats, 7 fish, 2
hamsters
RANK some
things as having
more of
something than
others (but NOT
QUANTIFY how
much of it they
have)
Interval
QUANTIFY how
much of
something there
is and a score of
zero means the
absence of the
thing being
measured
PSY103 Quantitative Methods in Psychology
Study Session 7
Frequency Distribution
Introduction
In trying to understand, interpret, predict, and modify behaviour,
psychologists engage in various forms of research that involve
data collection. Such data represent observable behaviours that
are recorded in numerical forms. This study session will expose
you to how to summarize data into manageable forms, using
frequency distributions.
When you have studied this session, you should be able to:
Demonstrate how data can be put in a simple
frequency distribution pattern. (SAQ7.1)
Learning construct class intervals, and enhance the use of tally
Outcomes marks with subsequent frequency values. (SAQ7.2)
determine exact limits and mid-point values in
grouped frequency distribution. (SAQ7.2)
between 0 and 10. Look at the scores for the group of 50 trainees
in Table 6.1.
1 | 1
0 0
Σ =N=
50
7.1.1 Steps in Constructing a Regular Frequency
1 List every raw score value in the first column of a table and
denote this by the symbol [X], with the highest score at the
top and the lowest score at the bottom. Thus, 10 is at the top
while 0 is at the bottom.
2 Run through the raw data either horizontally or vertically,
recording each observed score against identical score in the
first column with a tally mark. Show the tally marks in the
second column. After every four tallies, the fifth tally is
used to cross the previous four (it is a rule). It is aimed at
easing the recording of frequencies.
3 When the tally is completed, the tally marks for each score
in the first column are added up to find the frequency (f) and
recorded in the third column. The frequency of each score
shows the number of times a given score was obtained. You
can easily read from Table 7.2 that two students had a score
of 10, five students had a score of 9, seven students had a
score of 8, etc. Cross check with Table 7.1
4 Add up the frequencies and place the total at the bottom of
the frequency column. The Greek sign Σ (Sigma) means the
―sum of‖ and Σf means ―the sum of frequencies‖ is used to
show this. It means the total number of scores in the sample.
Σf and N must be equal, they are therefore used to check
error in tallying. Thus, if Σf and N are not equal, the tallying
exercise should be repeated. Even when they are equal, it is
still advisable to repeat the exercise to check for error of
placing one or more tally marks wrongly. If the first tallying
was done vertically, the second may be done horizontally to
check for such error.
Study Session 7 Frequency Distribution
Step 2: List the heights (in order of size) in the first column,
the tally in the second column, and the frequency in the third
column.
63 |||| 2
64 || 4
65 ||| Fill the
66 remaining
columns
67 accordingly
68
69
+1
Where C.I. = Class interval
H = Highest score
L = Lowest Score
I = Class Interval size chosen
H –L = Range
For the case under study we have:
Assessment
Now that you have completed this study session, you can
assess how well you have achieved its Learning Outcomes by
answering these questions.
Assessment
SAQ 7.1 (tests Learning Outcome 7.1)
A security unit of a psychological centre measured the speed
of 25 cars that entered into its premises. The resulting speeds
were:
29, 23, 30, 30, 27, 24, 30, 25, 23, 28, 25, 24, 28, 30, 23, 30,
27, 25, 29, 24, 23, 26, 30, 28, 25
Prepare a frequency distribution for this data.
Assignment
Study Session 8
Graphic Representation of
Frequency Distribution
Introduction
So far, Study Session 7 has shown how data could be collected
and organized statistically. As earlier explained, statistics also
involve the presentation and interpretation of data. This Study
Session will show how facts of frequency distribution can be
presented clearly in some graphic ways. Such presentations help
to interpret the statistical data, by translating numerical
information, which are sometimes difficult to comprehend into
pictorial forms that are more readily understandable. Graphic
presentations, thus give better picture of a frequency distribution
by setting out its general contour and showing more precisely the
number of cases in each interval.
When you have studied this session, you should be able to:
present frequency distributions in polygon, histogram,
Learning bar chart, and relative bar chart forms. (SAQ 8.1)
Outcomes
In this Study Session, you have seen that there are three basic
types of graphs that we use for most data: (1) frequency
polygon, (2) bar chart(graphs), and (3) histograms. The simple
Summary frequency polygons are adequate for continuous data where
scores can emerge between score values. The names of the
last two are a bit misleading because both are created using
bars. The only difference between a bar graph and a histogram
is that in a bar graph the bars do not touch while the bars do
touch in a histogram. In general, bar graphs are used when the
data are discrete or qualitative. The space between the bars of
a bar graph emphasize that there are no possible values
between any two categories. For example, when graphing the
number of children in a family, a bar graph is appropriate
because there is no possible value between any two categories
(e.g., 1 and 2 children). When the data are continuous, we use
a histogram. The bars touch in a histogram to indicate that
there are possible values between any two categories. For
example, if we were graphing time to complete a test, the bars
would touch to indicate that there are possible values between
any two times (e.g., 27 and 28 minutes).
Assessment
Now that you have completed this study session, you can
assess how well you have achieved its Learning Outcomes by
answering these questions.
Self Assessment
SAQ 8.1 (tests Learning Outcome 8.1)
Draw a bar graph of frequency distribution car speed in
SAQ7.1
Study Session 8 Graphic Representation of Frequency Distribution
Now that you have completed this study session, attempt the
following assignment (a tutor marked assessment).
The table below shows the population of undergraduate
students in the Faculty of the Social Sciences at the University
Assignment of Ibadan for the 2005/2006 academic session.
Study Session 9
Diagrammatic Representation of
Frequency Distribution
Introduction
Statistics involve, among others, the presentation and
interpretation of data in some diagrammatic ways. This Study
Session will discuss the representation of frequency distribution
using pie chart, pictogram and stem-and-leaf display.
When you have studied this session, you should be able to:
present raw data in the form of pie chart, pictogram,
Learning and stem-and-leaf display. (SAQ 9.1)
Outcomes
Works 98
Total 498
For Accounts, it is
For Production, it is
For Personnel, it is
For Works, it is
Angle of sector
For Marketing it is
For Accounts, it is
For Production, is
For Personnel, it is
For Works, it is
Note that the angles must sum up to 360o. See Table 9.3
Table 9.3 Pie Chart Distribution of the Employees in a Company in
Table 9.1
Department Student Relative Angle of
Population Frequencies Sector
Marketing 153 30.7% 110.60o
Accounts 60 12.0% 43.37o
Production 102 20.5% 73.74o
Personnel 85 17.1% 61.45o
Works 98 19.7% 70.84o
Total 498 100% 360o
3. Draw the Pie Chart: This can be done neatly with a pair of
compass and a protector. Using a pair of compass, measure
any length and draw a circle, showing the circumference.
From the center of the circle, draw a straight line to the
Study Session 9 Diagrammatic Representation of Frequency Distribution
In this Study Session, you have learnt that the pie chart is a
circular diagram in which classes or attributes are represented
by sectors. A pictogram is a graphic way of presenting data
Summary with emphasis on the aesthetic aspect of the display and not
accuracy. The stem-and-leaf display is a technique for
summarizing a set of data by combining the features of the
frequency distribution and the histogram.
Assessment
Now that you have completed this study session, you can
assess how well you have achieved its Learning Outcomes by
answering these questions. Write your answers in your Study
Self Assessment Diary and discuss them with your Tutor at the next Study
Support Meeting. You can check your answers with the Notes
on the Self-Assessment Questions at the end of this Manual
SAQ 9.1 (tests Learning Outcome 9.1)
The table below shows the monthly spending of Mrs. Halima
ITEM AMOUNT SPENT (N:K)
Food 7000
Transportation 1250
Accommodation 6000
Savings 1500
Miscellaneous 1000.00
Study Session 10
Cumulative Distributions
Introduction
In the simple frequency distribution, you can easily determine the
number of cases falling within a class interval by adding up
frequencies in that interval. But when additional information is
required, like the number of cases occurring above or below a
given interval, then cumulative frequency distribution becomes
necessary. This Study Session will expose you to how to use the
cumulative frequency distribution which is an alternative method
that is used when information is required in form of ―more than‖
or ―less than‖.
When you have studied this session, you should be able to:
infer from frequency distributions, the number of cases
falling at, and above or less than, a particular point.
Learning plot graphs showing ―more than‖ and ―less than‖
Outcomes cumulative frequency distributions.
73-77 2 2 50
68-72 3 5 48
63-67 4 9 45
58-62 5 14 41
53-57 5 19 36
48-52 7 26 31
43-47 6 32 24
38-42 4 36 18
33-37 4 40 14
28-32 3 43 10
23-27 4 47 7
18-22 3 50 3
Σ f = N = 50
***Please the title for all Figures are supposed to be below the
figure not above. Only Tables should have their titles above.
Please take not for all corrections
When all the cumulative frequencies have been plotted, the points
are joined by straight lines. To bring the curve down to the base
line, the cumulative frequency of 0 is plotted at the upper exact
limit of the interval 73-77, which is 77.5. This means that there
are no employees with performance scores of ―more than‖ 77.5.
10.2.2 Graphical Presentation of ―Less Than‖ Cumulative
Distribution.
In the ―less than‖ cumulative frequency, the graph is plotted
between the cumulated frequencies and the upper exact limit of
the interval. The graph is brought down to the base by placing the
cumulative frequency of 0 at the lower exact limit of the bottom
interval. Thus, there is no employee who performed less than
17.5.
From the graph in Figure 10.2, you can now easily read the
number of employees who performed less than any given score.
So also, you can read the number of employees who performed
more than any given score, with the ―more than‖ cumulative
distribution.
PSY103 Quantitative Methods in Psychology
Assessment
Now that you have completed this study session, you can
assess how well you have achieved its Learning Outcomes by
answering these questions. Write your answers in your Study
Self Assessment Diary and discuss them with your Tutor at the next Study
Support Meeting. You can check your answers with the Notes
on the Self-Assessment Questions at the end of this Manual
SAQ 10.1 (tests Learning Outcome 10.1)
Cumulative frequency distribution is that which helps to calculate
the number of people falling at a particular score:
(a) True
(b) False.
Study Session 11
Introduction
The frequency distribution with its graphic representations
primarily shows the organization of data into a manageable form.
They are the preliminary steps toward quantitative treatment of
data. The method of statistically treating the distribution gives
numerical descriptions which are the basic features of the
distribution. These further enable concise and definite comparison
of the features of one distribution with others. This study session
will attempt to expose you to measures of central tendency, which
is one of the types of numerical descriptions of frequency
distributions.
When you have studied this session, you should be able to:
compute and interpret the arithmetic mean or average
score for both ungrouped and grouped data.
Learning
Outcomes
???
ΣX = 404
N=7
PSY103 Quantitative Methods in Psychology
Assessment
1.
Assignment
The Table 11.3 are data of time spent by students reading for an
examination. Study the Table closely and answer the questions
that followTable 11.3: Showing Raw Data of Time Spent by
Students Reading for an Examination
6 2 2 6 5 7 4 9 8 5
4 9 5 6 7 5 3 5 5 6
3 7 8 4 3 7 6 7 6 6
8 6 5 4 10 5 3 9 4 8
(a) Using the date in Table 11.3, compute the average time spent
by the students reading(b) Using a grouped frequency
distribution, compute the average time spent by the students
reading
Post your answers on Study Session 11 assignment page at UI
Mobile Class.
PSY103 Quantitative Methods in Psychology
Study Session 12
Introduction
This study session will expose you to how scores are arranged.
The median is the mid-point of a set of scores. It is the point on a
continuum of scores above which fall exactly one half of the cases
and below which fall the other half. When scores are arranged in
order of magnitude, either in ascending or descending order, the
median has the same number of scores above and below it.
When you have studied this session, you should be able to:
locate the median position of a set of ungrouped data,
as well as calculate the median value of grouped
Learning frequency distribution, with and without class intervals.
Outcomes determine the median value of a distribution, using the
distribution curve, called OGIVE.
have 7, 8, 8, (9, 9), 11, 13, 15. The 4th and 5th observations are 9
and 9 respectively. Md is ½ (9 + 9) = ½ (18) = 9.
Md
Where LRL = lower real limit of median interval
N/2 = one half of the total number of cases
SFB = sum of all the frequencies up to, but not
including the median interval
F = the frequency within the median interval
I = class interval size
12.2.1 The Median of Grouped Data with Regular Frequency
Distribution
The computational procedures shall be explained, using the data
in Table 11.1, which is now shown in Table 12.1 for the median
calculation.
Table 12.1 Regular Frequency Distribution of Accidents recorded in
a Company
No. of No. of Employees
Accidents (f)
(X)
9 2
8 2
7 6
6 5 5 cases down here
5 3 Median interval
4 4 13 cases up here
3 5
PSY103 Quantitative Methods in Psychology
2 4
1 0
Σ f = N = 31
F =21
I =10
Theoretical Explanations
In Table 11.2, the median interval is the interval 40 – 49 since the
50th case resides within the interval. Remember, the 50th case is
the median point for the distribution as N/2 i.e. 100/2 is 50.
Thirty-one cases are counted up to the top of the interval 30-39
(less than cumulative frequency). The next interval 40-49, which
is the median interval contains 21 cases, suggesting that the 50th
case lies somewhere within this interval because adding 21 to 31
gives 52; this is greater than the 50th case desired. The problem
now is how to determine where the 50th case lies within the
median interval?
For the sake of interpolation, we assume that the 21 cases to the
median interval are evenly distributed over the whole range of the
interval with exact limits as 39.5 and 49.5. Actually, we require
19 cases of the 21 cases in the median interval to be added to the
sum of the preceding cases (i.e. 31) to make the desired 50 cases.
Thus, we need a fraction of 19/21 of the median interval. Since
the class width of the median interval is 10, it implies we desire
19/21 of 10, which is 9.05 units of the interval to correspond with
the 50th case desired. Adding the 9.05 to the exact lower limit of
the interval (39.5), we now have 39.5 + 9.05, which is 48.6 as the
median value. So 48.6 corresponds with the 50th case in the
distribution. You can as well observe that 48.6 falls within the
median interval.
Table 12.2 Calculation of the Median of Performance Scores for 100
Employees
Class Interval F Cf
90-99 2 100
80-89 4 98
70-79 3 94
60-69 15 91
50-59 24 76 48 cases down
here
40-49 21 52 Median
Interval
PSY103 Quantitative Methods in Psychology
12.3 Determining the Median using the Cumulative Frequency (Or OGIVE Curve)
To find the median of grouped data, the cumulative frequency
distribution curve can as well be used. From the curve, we can
easily read off the median, having determined the median point as
N/2 on the Y-axis for the cumulative frequency.
Figure 12.1 is the OGIVE of data in Table 12.2, employing the
―less than‖ cumulative frequency distribution. The median is the
dividing point of the desired point, that is N/2, which is 100/2 =
50. Read off 50 on the Y-axis to meet the curve at N. From N,
draw another dotted line perpendicular to the Y-axis to meet the
X-axis at P. The corresponding value of P on the X-axis is the
median of the data.
Assessment
Now that you have completed this study session, attempt the
following assignment (a tutor marked assessment).
1. Ten employees in an organization reported the years
they have spent on the job as follows:
Assignment
14, 20, 16, 14, 12, 10, 18, 16, 11, and 12
What is the median year of the employees on the job?
Study Session 13 Measures of Central Tendency: The Mode
Study Session 13
Introduction
This Study Session aims to provide explanations on the mode,
which is the most frequent score in a distribution of scores. The
mode is the value with the highest frequency.
When you have studied this session, you should be able to:
compute the modal value for ungrouped and grouped
frequency distributions with and without class
Learning
Outcomes intervals.
Given a set of scores like these in Table 13.1, the mode is 5. The
highest frequency is 15, while the corresponding score is 5,
indicating that the score 5 occurred 15 times in the set of data.
Thus, 5 is the mode, being the most frequent score.
Table 13.1 Determining the Mode of a Distribution
Score 0 1 2 4 5 7 8 9
Frequency 7 3 11 7 15 6 4 5
1. If there are more than one class intervals separating the two
class intervals, like distribution A in Table 13.2, then we
say that the distribution has two modes or it is bimodal.
The modes will be the two mid-points of the respective
class intervals. Thus, we have 9.5 and 3.5 as the modes for
distribution A in Table 13.2.
2. If the two intervals are separated by only one intervening
interval, like distribution B in Table 13.2, then it is
possible that the distribution is actually unimodal. It is
more certain to be unimodal when the intervening interval
has a relatively high frequency. In such a case, the crude
mode is not adequate as it will not be possible to decide
what the crude mode is.
3. When the two class intervals with the highest frequency
occupy adjacent positions like distribution C in Table 13.2,
then the crude mode is decided as the dividing point
between the two intervals. The dividing point is simply the
upper real limit of the lower class interval, which is
equivalent to the lower real limit of the higher of the two
class intervals. For distribution C in Table 13.2, the crude
mode is 6.5. Such distribution is invariably unimodal.
Table 13.2 Illustrations of when the Crude Mode is Adequate
Distributions (Frequency)
Class Interval
A B C
11 – 12 2 2 2
9 – 10 7 7 6
7–8 6 6 7
5–6 2 7 7
3–4 7 2 2
1–2 3 3 3
Bimodal Unimodal
= 49.5 + 3/12 x 10
= 49.5 + 2.5 = 52
The crude mode is 54.5 while the interpolated mode is 52, which
is a ―pull away‖ towards the preceding interval. The reason is that
the preceding interval has a frequency of 21, while the following
interval has a frequency of 15; hence, the pull towards the
preceding interval. If the frequency of the following interval was
larger, then the pull would have been directed towards the
following interval, making the interpolated mode above 54.5 but
less than 59. Thus, the interpolation of the mode within the modal
class interval improves the estimate of the mode by allowing the
PSY103 Quantitative Methods in Psychology
The Mean
1. The arithmetic mean is often used when the distribution is
reasonably symmetrical. This is because the mean takes
into account all measures in a distribution. The extremely
small values and the extremely large ones, as well as those
near the centre of the distribution, influence the mean.
When, however, there are extreme values on one side of
the distribution that are not balanced by extreme values on
the other side, the mean is greatly influenced. This results
in a misleading picture of the location of the measure of
central tendency. Consequently, the mean is not adequate
for a distribution that is highly skewed.
2. For a normal distribution, the sum of differences (or
deviations) of all scores from the mean is zero. That is, X –
Mean = 0. This implies that the mean balances the sums of
the positive and negative deviations. This will be expanded
later.
3. The mean is the most reliable and stable of the three
measures of central tendency. This is because when
different samples are drawn from the same population and
their means computed on the same measure, there will be
little or no difference in the means. Also, errors of
measurement tend to neutralize one another around the
arithmetic mean. If any error exists in the use of a mean,
the error is considerably smaller than the error of a single
measure in the distribution.
4. The mean is useful in further statistical computation, while
other measures of central tendency are usually not. For the
characteristics of the mean, it is used in many of the
procedures of inferential statistics.
The Median
1. The median is a positional measure of central tendency. Its
position is determined directly by the number of values in a
distribution, not the magnitude of the values. The
magnitude of values only indirectly affects the median by
determining the serial ordering before decision on what the
median is.
2. Though all measures contribute to the calculation of the
median, just like the mean, the median is less affected by
the extreme of the distribution than is the arithmetic mean.
For example, the median of the scores 3, 5, 7, 9, and 9 is 7
while the arithmetic mean is 6.6. If the extreme values are
changed and we now have 4, 6, 7, 12, and 13, the median is
PSY103 Quantitative Methods in Psychology
till 7 but the arithmetic mean is now 8.4. If the scores are 4,
6, 7, 9, 9, 12, and 13 then the median becomes 9. Thus, the
median is mainly influenced by the number, rather than the
size of the extreme values in a distribution.
3. The median is an easy measure of central tendency to
compute.
The Mode
1. It is the quickest estimate of central tendency available.
2. It is the most typical measure in that it tells the most
occurring in a set of scores.
Assessment
Study Session 14
Measures of Variability
Introduction
A measure of central tendency tells much about a distribution but
does not give a complete picture of the distribution. When two
distributions are to be compared, if it is based solely on averages,
quite wrong conclusions may be drawn. Other descriptive
statistical measurements need to accompany averages to amplify
the description of a set of data, hence this Study Session focus on
measures of variability are essential.
When you have studied this session, you should be able to:
differentiate between measures of central tendency and
measures of dispersion or variability.
Learning compute and interpret measures of dispersion like
Outcomes range, mean absolute deviation, standard deviation, and
variance . (SAQ 14.1)
3 3 2
3 3 1
3 2 1
3 2 0
The arithmetic mean of each of the three distributions is the same,
i.e. 5, while the median and mode is 3 and it is the same for the
three distributions. If we are to compare the three distributions on
the basis of the measures of central tendency, then we may
conclude that the distributions are the same. This may be rather
erroneous because a close observation of the scores in each
distribution shows that they are dispersed differently. For
instance, the lowest score in distribution A is 3 while the largest is
8 and the scores do not vary much from one another. For
distribution B, the lowest is 2 while the largest is 10 and the
scores are more spread out than distribution A. Distribution C has
0 as the lowest score and 20 as the largest. The scores in
distribution C are more dispersed than the other two distributions.
A measure of dispersion or variability is needed for a more
accurate comparison to be made among the distributions. Thus,
the central locations of two distributions may be the same but the
variability may differ. Also, the variability may be the same but
the central locations may differ. Locations and variability are
actually independent. While variability refers to the distance
between each score and any other and is measured with the
measures of variability, the measures of central tendency show
locations in a set of data. A measure of variability show,
numerically, the degree to which scores spread around the
average. It is a measure which tells whether scores cluster closely
around the mean or whether they scatter widely.
Measures of variability are useful in many areas of life, although
it is not frequently reported because it is not familiar to many as
the measures of central tendency. A measure of variability tells us
how representative the average is. If the numerical value of the
measure is small, it means that the individual scores are close to
the average. If the variation score is large, the mean can be used
with less assurance because the scores are rather far from the
mean.
between the highest value, H, and the lowest value, L. This can be
symbolized as:
R=H–L
Where R = range
H = highest value
L = lowest value
For grouped data, the range is calculated by subtracting the lower
limit of the bottom interval from the upper limit of the top
interval. Thus,
R = Ult – Llb
Where R = range
Ult = upper limit of top interval
llb = lower limit of bottom interval
The range is, however, the most unreliable of the measures of
variability. This is because it is determined by only two measures,
with all the other individual values in between having no effect in
the calculation. There is often a gap between those extreme values
used in determining the range and the next highest or lowest value
in a distribution. If these extreme cases are absent from the
distribution then there would be a significant difference in the
calculated range. For instance, consider these two distributions.
Distribution A: 20 16 16 15 15 15 15 14 14 10
Distribution B: 20 19 19 18 17 16 15 14 11 10
The range for the two distributions is 20 - 10, which is 10. An
observation of the values in distribution A shows that the scores
are relatively clustered near the middle of the distribution while
those of distribution B are more spread out. The range is therefore
a weak measure of variability of distribution A. For instance,
when the two extreme scores are eliminated from the two
distributions, the range for distribution A becomes 16 - 14, which
is 2; this is rather far from the earlier range of 10. For distribution
B, the range becomes 19 - 11, which is 8; it is still close to the
initial range of 10. Since this type of change occurs frequently,
the range is considered as a crude measure of variability, just like
the mode, and it is avoided as much as possible in the behavioural
sciences.
Generally, the range is more unreliable with a small sample size
than a large sample size. This is because with a small sample size
the extreme values may be marked by one case. But when the
PSY103 Quantitative Methods in Psychology
sample size is large, there may be more than one case of each of
the extreme values. Thus, the probability of the range being
altered as a result of missing scores is higher for a smaller sample
than for a larger sample.
= 16.2
= Individual Score
Deviation score =
Where X = individual score in the distribution
to replace X
Then we will always have a zero score, which will make nonsense
of the whole exercise. To overcome this problem, we have to take
the absolute value of each deviation by ignoring the sign, and then
utilizing its numerical value for the computation of the mean
absolute deviation (MAD). Though the MAD is a descriptive
measure of variability, but it is usually rejected for its inadequacy
in further statistical analysis. This is because absolute values
which are incorporated in MAD are unsuitable for use in
inferential statistical analysis.
The measure which takes care of this problem without using
absolute values, and is useful for further statistical analysis is that
which requires the squaring of each of the deviations before
taking the average of the squared deviations. By squaring the
deviation, the negative signs are taken care of and all the squared
deviations become positive. The sum of the squared deviations
from the mean, that is, Σ(X – )2 is called the sum of squares and
is symbolized by SS. The average of the sum of squares is a
measure of variability, called variance, which is symbolized by δ2.
These formulae are called definition formula. They are used for
manual calculation.
The range is the difference between the highest value and the
lowest value in a distribution. The mean absolute deviation is the
deviation of each measure in a distribution from a measure of
central tendency. The standard deviation is the positive square
root of the average sum of square deviations from the mean, while
the square of this value is the variance. Thus, the average of the
sum of squares is variance.
Study Session 14 Measures of Variability
Assessment
SAQ 14.1
Given the following frequency distribution, find the standard
Assessment deviation of the data.
X 6 7 8 9
F 2 3 3 2
Number of 9 8 7 6 5 4 3 2
Accidents
Drivers involved 2 2 6 5 3 4 5 4
Study Session 15
Transformed Scores
Introduction
If you are informed that you have obtained 75% in PSY 103
examination, you will definitely need additional information in
order to determine how well you performed, compared with other
students who took the examination. If the examination was
generally easy for most students and there were many high scores,
then your score of 75% may be an average or even below average
score. If, on the other hand, the examination was a little bit
difficult for the whole class, then your score may be among the
highest or even the highest. One way of getting this additional
information is to transform the original score to show at a glance
how well you performed in comparison to other students in the
course. Thus, this Study Session addresses the different ways of
transforming scores to allow for relative comparison of an
individual‘s scores across different tasks, as well as relative
comparisons of individuals‘ scores on the same task.
When you have studied this session, you should be able to:
compute transformed scores.
interpret percentile, standard score (Z score), and T
Learning score.
Outcomes
15.1 Percentiles
The percentile rank of a score is a single number which gives the
percentage of cases in the specific reference group, scoring at or
below that score. If, for example, your raw score in an
examination is 75%, and it corresponds to a percentile rank of 65,
it means that 65% of the class obtained equal or lower scores than
you did, while 35% of the class received higher scores. A
percentile is, thus, the score at or below which a given percent of
the case lie. The percentile shows directly how an individual score
compares to the scores of a specific group. It is important to keep
in mind that a percentile rank cannot be correctly interpreted,
unless the reference group is taken into consideration.
Study Session 15 Transformed Scores
Step 1: Locate the class interval wherein the raw score falls; call
this the ―critical interval‖.
Step 2: Group the frequencies (F) into three categories,
embracing those with scores higher than the critical interval, those
corresponding to all scores in the interval, and those
corresponding to all scores lower than the critical interval. These
are 6, 6, and 75 respectively.
Step 3: Each frequency is then converted to a percent by dividing
by N; in this case 87. We will denote the percent of people
scoring in interval higher than the critical interval by H% (for
higher), the percent of people scoring in the critical interval by I%
PSY103 Quantitative Methods in Psychology
(for in), and the percent of people scoring lower than the critical
interval by L% (for lower)
Step 4: It is clear at a glance that the score of 43 is better than at
least 86.2% of the scores, i.e., those below the critical interval.
Thus, the percentile rank must be at least 86.2%. It is also clear
that 6.9% of the scores are better than 43, i.e. the ones above the
critical interval. However, it is not evident whether the raw score
of 43 is higher than all the scores that fall within the critical
interval, or it is lower than, since the interval carries 6.9% of the
entire scores. The solution is to look at the score in comparison to
the size of the interval; the higher the score in relation to the
critical interval the more people in that interval is assumed to
have been outscored.
Step 5: To determine the standing of score 43 in the critical
interval, first ascertain the lower real limit of the interval, which is
41.5. Subtract this from the raw score of 43 (43 – 41.5) = 1.5.
Since the interval size is 3, this distance is expressed as a fraction,
and is equal to 1.5 points/3 points, = 0.5 of the interval. Thus,
score 43 is 0.5 of the 6.9% of students in the critical interval. This
gives a value of 3.45, which should be added to 86.2% (i.e. L %).
The final result is 89.65%; implying that 89.65% of those that
took the examination had 43 or less, which is the percentile rank.
The percentile rank is determined from the formula:
pN = .32 x 87 = 27.84
Step 3: The 27.8th case must have a score of at least 26.5; the
lower real limit of the interval in which it appears. However, the
critical value covers three score points (26.5 – 29.5); what point
value corresponds to the 32nd percentile?
There are 23 cases below the critical interval, so that the 27.84th
case needs 27.84 – 23, i.e. 4.84 cases up in the interval. The total
number of cases in the interval is 9, so this distance expressed as a
fraction is 4.84 cases /9 cases = .54. Therefore, in addition to the
lower real limit of 26.5, .54 of the three points included in the
critical interval must be added in order to determine the point
corresponding to the 27.84th case. The cutting score is therefore
equal to 26.5 + (.54 x 3) = 26.5 + 1.62 = 28.12.
A general formula to obtain this is:
Where
Score p = Score corresponding to the pth percentile
LRL – lower real limit of critical interval
P – specified percentile
N – total number of cases
SFB – sum of frequencies below critical interval
F – frequency within critical interval
H – interval size
= mean score
= standard deviation
Converting each of the original test scores to Z scores yield:
PSY103 Quantitative Methods in Psychology
X 75 65 55
80 60 50
10 15 5
The standard scores show at a glance that the student was half
standard deviation below the mean in PSY 101, less than half
standard deviation above the mean in PSY 102, and one standard
deviation above the mean in PSY 103. When raw scores are
transformed into Z scores, the shape of the distribution remains
the same. The Z scores give an accurate picture of the standing of
each score relative to the reference group, no matter where the
original scores are measured from or what scale is used. Standard
scores are used extensively in the behavioural sciences.
15.2.2 T Scores
Standard scores are a little bit difficult to explain to one not well
versed in statistics. Since behavioural scientists try to report test
scores to people who are not statistically sophisticated, several
alternatives to Z scores have been developed. One of such
alternative, called T scores, is defined as a set of scores with a
mean of 50 and a standard deviation of 10. The T scores are
obtained from this formula:
T = 10Z + 50
For example, the score of 55 in the example is converted to a Z
score of +1.00 by using the Z formula. Then T is equal to (10)
(+1.00) + 50 = 60.
Since the mean of T scores is 50 it can still be seen at a glance
whether a score is above average (it will be greater than 50) or
below average (it will be less than 50). Also one can tell how
many standard deviations above or below average a score is. For
example, a score of 40 is exactly one standard deviation below
average (i.e. a Z score of -1.00) since the standard deviation of T
scores is 10. But the T score of 60 confirms the Z score of +1.00
which is interpreted as one standard deviation above the mean.
Study Session 15 Transformed Scores
Assessment
References
Feedback on SAQs
SAQ 4.2 The correct answer is B because the nth term is obtained by
dividing the sample frame or population (600) by the desired
sample size (200), giving 30.
If you chose A, you would be wrong because you may have
assumed that the nth term was the sample size.
If you chose C, you would be wrong. You may have taken the
sample frame or population as the nth term.
If you chose D, you would be wrong. You may have subtracted
the sample size from the sample frame or population.
SAQ 5.1 We do not know what you have provided in your table, but it may
be filled like as shown below:
23 //// 4
24 /// 3
25 //// 4
26 / 1
27 // 2
28 /// 3
29 // 2
30 //// / 6
PSY103 Quantitative Methods in Psychology
SAQ 8.1 Pie chart showing the monthly spending of Mrs Halima
SAQ 9.1 The answer is (b). The cumulative frequency distribution enables
the calculation of the number of cases or people falling at, below,
or above given scores in the distribution.
SAQ 13.1 Variance formula is:
Study Session 16
The Normal Distribution
Introduction
This study session presents the normal distribution and the skewed
distributions. It explains a normal distribution and its graphical
presentation. The empirical rule and parameter of the normal
distribution are presented, in terms of the mean and standard deviation.
The properties of the normal distribution are compared to those of the
skewed distributions and the locations of the three measures of central
tendency on the normal curve, positively skewed and negatively
skewed distributions are given
you‘ll find extreme values far from the peak on the low side more
frequently than the high side.
The mean, median, and mode are all equal in the normal distribution
and other symmetric distributions.
Questions
4. Present a graph of a normal curve showing the locations of the mean, mode,
and median
Probability sampling techniques involve some form of random selection and provide each member of the population with an equal chance of being selected, which supports scientific objectivity and generalizability of results . On the other hand, non-probability sampling techniques do not use random selection, which can lead to biased samples as not all members have equal chances of being selected, thus limiting the scientific objectivity .
Bias can be introduced in research using non-probability sampling techniques because these methods do not ensure every population element has a chance of selection, leading to samples that do not represent the population adequately. Purposive sampling relies on researcher judgment for selection, potentially introducing researcher bias, while accidental sampling may result in convenience-biased samples that are not indicative of the wider population .
Proper construction of class intervals in a grouped frequency distribution ensures accuracy by ensuring intervals are mutually exclusive and collectively exhaustive, preventing overlaps and covering all observations . Equal interval sizes facilitate consistent comparison and summary, and ensuring class intervals accommodate the range of data avoids misrepresentation . This design simplifies the visualization of data distributions and enhances the interpretability of statistical analyses .
Cumulative frequency distributions aggregate data to show how many observations fall above or below a certain point. A 'more than' distribution begins with the highest interval and descends, showing cumulative totals for values above the interval . Conversely, a 'less than' distribution starts with the lowest interval and ascends, indicating cumulative totals for values below the interval . They are useful for understanding the spread and concentration of data within a dataset .
Mean absolute deviation provides insights into variability by averaging the absolute differences between each data point and the mean, indicating average dispersion without removing sign direction, offering clarity in understanding spread around the mean . Unlike MAD, standard deviation measures variability by considering the spread of data as it squares the deviations, placing more weight on larger deviations, which provides a detailed understanding of data spread, especially for normally distributed data .
Determining the nth term in systematic sampling involves dividing the population size by the number of desired participants to calculate the interval for selecting each nth participant . This process contributes to sample accuracy by ensuring equal representation across the population, reducing potential bias from random chance and providing an efficient means of creating evenly spaced samples .
Snowball sampling operates by using existing study participants to recruit future participants through referrals, which is advantageous in researching unknown or rare populations by accessing subjects that may not be reachable through other sampling methods . However, it carries potential limitations such as recruitment bias, as the sample may not be representative of the wider population due to dependence on social networks of initial participants .
The empirical rule states that in normal distributions, approximately 68% of data falls within one standard deviation from the mean, 95% within two, and 99.7% within three standard deviations. The importance of standard deviation in this context lies in its role as a measure of variability, defining the spread around the mean; it provides insights into the dispersion patterns typical of a normal distribution and helps identify outlier data points .
Scaling in research translates qualitative observations into quantitative data, allowing statistical analysis. Nominal scales categorize data without any order (e.g., gender); ordinal scales rank data with order but without fixed intervals (e.g., class rankings); interval scales position data with equal intervals but no true zero (e.g., temperature); ratio scales, like interval scales, include a true zero point, permitting meaningful ratios (e.g., weight). Each type of scale offers different potential for statistical operations based on their inherent properties .
Cluster sampling is considered more objective because it involves random sampling within pre-determined clusters, thus allowing for heterogeneity representation of the population and reducing overall sampling bias . Although simple random sampling also provides equal selection chances, it may not always capture population diversity. Non-probability techniques like purposive sampling do not use randomization, allowing potential bias and lack of scientific rigor in sample selection .