Frequency
Distributions
and Percentiles
Course: Psychological Statistics
Instructor: Marlon M. Dinoy
Review parameter
Terms: statistics statistic
population sampling error
sample descriptive statistics
random sample inferential statistics
variable constructs
data operational definition
data set discrete variables
datum continuous variables
Review
Terms: nominal scale independent variables
ordinal scale dependent variables
interval scale
ratio scale
descriptive research
correlational research
experimental research
nonexperimental research
Additional Terms:
Measurement is the use of a rule to assign a number to a specific
observation of a variable.
Validity refers to how well the measurement rule actually measures
the variable under consideration as opposed to some other variable.
Reliability is an index of how consistently the rule assigns the same
number to the same observation.
Additional Terms
Objectives:
At the end of the lesson, the students are able to:
explain the importance of frequency distribution;
describe some elements of frequency distribution and how they
are constructed and graphed;
describe distribution; and
define and determine percentile ranks and percentiles.
Student Loan Debt
From: Office of Federal Student Aid
Frequency distributions
Frequency distributions summarize data by categorizing and
counting.
A frequency distribution is a tabulation of the number (count) of
occurrences of each score value (category).
Frequency distributions
Frequency distributions
Using frequency distribution has benefits.
Frequency distributions
Relative frequency of a score value is the proportion of
observations in the distribution at that score value. A relative
frequency distribution is a listing of the relative frequencies of
each score value.
Frequency distributions
Advantage of relative frequency:
It combines information about frequency with information about
the number of measurements.
It is a good guess for the relative frequency of that score value in
the population from which the random sample was selected.
Frequency distributions
Frequency distributions
A cumulative frequency distribution is a tabulation of the
frequency of all measurements at or smaller than a given score
value.
A cumulative relative frequency distribution is a tabulation of the
relative frequencies of all measurements at or below a given score
value.
Not appropriate for nominal data
Useful in calculating percentiles
Frequency distributions
Grouped Frequency distributions
What if your data is more variability?
Like this...
Grouped Frequency distributions
Grouped Frequency distributions
A class interval is a range of score values. A grouped frequency
distribution is a tabulation of the number of measurements in each
class interval.
Grouped Frequency distributions
Some terms related to:
Class limits are the end numbers used to define the class
intervals. The lower class limit is the lower end number while
the upper class limit is the upper end number.
Open class interval is a class interval with no lower class limit
or no upper class limit.
Class boundaries are the true class limits. It has lower class
boundary and upper class boundary.
Class size is the size of the class interval.
Class mark is the midpoint of a class interval.
Grouped Frequency distributions
Grouped Frequency distributions
Guidelines for grouped frequency intervals:
There should be between 8 and 15 intervals.
Use convenient class interval sizes, like 2, 3, 5, or multiples of 5.
Start the first interval at or below your lowest score.
Graphing Frequency Distribution
The abscissa or x-axis is the horizontal axis. For frequency and
relative frequency distributions, the abscissa is marked in units of
the variable being measured, and it is labelled with the variable’s
name. The ordinate or y-axis is marked in units of frequency or
relative frequency, and so labelled.
A relative frequency histogram uses the heights of bars to
represent relative frequencies of score values (or class intervals).
Graphing Frequency distributions
Describing Distribution
Distributions can be described and characterized in three major
ways:
shape,
central tendency, and
variability.
Describing Distribution
Symmetry and skew are ways in which the shape of a distribution
can be characterized.
A symmetric distribution can be divided into two halves that are
mirror images of each other.
A positively skewed distribution has score values with low
frequencies that trail off toward the positive numbers (to the right). A
negatively skewed distribution has score values with low
frequencies that trail off toward the negative numbers (to the left).
Describing Distribution
Modality refers to the number of modes, or peaks, in the
distributions. It is another aspect of the shape of a distribution.
A unimodal distribution has one peak. A bimodal distribution has
two peaks. A multimodal distribution has more than two peaks.
A distribution that is unimodal and symmetric is called a bell-shaped
distribution.
Describing Distribution
Describing Distribution
The central tendency of a distribution is the score value near the
center of the distribution. It is often a typical or representative score
value. e.g. mean, median, and mode
Describing Distribution
Variability is the
degree to which
the measurements
in a distribution
differ from one
another.
e.g. range,
standard deviation,
variance
Describing Distribution
Using the three major characteristics to compare and contrast
different distributions.
Percentiles
Percentiles define a measurement’s relative standing in the
distribution.
e.g. percent correct of your score in a Statistics test vs. percentage
of the number of individuals who got scores lower than your score
The percentile rank of a score value is the percent of
measurements in the distribution below that score value.
The Pth percentile is the score value with P% of the measurements
in the distribution below it.
Percentiles
About percentiles and percentile ranks:
They can be used if the data is not in nominal scale
They are only approximate.
Percentile ranks are an ordinal index of relative standing.
Percentile ranks can be interpreted only within the context of a
specific distribution.
They can be used to make boxplots.
Other terms related to percentiles:
The quartiles divide the ordered observations into 4 equal parts.
The deciles divides the ordered observations into 10 equal parts.
References
Glenberg, A., & Andrzejewski, M. (2007). Learning From Data: An
Introduction To Statistical Reasoning (3rd ed.). Routledge.
[Link]
Gravetter, F. and Wallnau, L. (2015) Statistics for the Behavioral
Sciences. 10th Edition, Cengage/Wadsworth, Boston.
Almeda, J. V., Capistrano, T. G., & Sarte, G. M. F. S. (2010).
Elementary statistics. University of the Philippines Press.
...