Unit II Descriptive Analytics
Unit II Descriptive Analytics
Data
A collection of actual observations or scores in a survey or an experiment is termed as data. Data can be
categorised into three types. They are qualitative data, ranked data, or quantitative data.
Qualitative Data
A set of observations where any single observation is a word, letter, or numerical code that represents a
class or category is termed as qualitative data. Qualitative data consist of words (Yes or No), letters (Y or
N), or numerical codes (0 or 1) that represent a class or category.
Ranked data
A set of observations where any single observation is a number that indicates relative standing is known
as ranked data. Ranked data consist of numbers (1st, 2nd, . . . 40th place) that represent relative standing
within a group.
Quantitative Data
A set of observations where any single observation is a number that represents an amount or a count is
known as quantitative data. Quantitative data consist of numbers (weights of 238, 170, …., 185 lbs) that
represent an amount or a count.
For example, the weights reported by 53 male students in Table 1.1 are quantitative data, since any
single observation, such as 160 lbs, represents an amount of weight. If the weights in Table 1.1 had been
replaced with ranks, beginning with a rank of 1 for the lightest weight of 133 lbs and ending with a rank
of 53 for the heaviest weight of 245 lbs, these numbers would have been ranked data, since any single
observation represents not an amount, but only relative standing within the group of 53 students.
Finally, the Y and N replies of students in Table 1.2 are qualitative data, since any single observation is a
letter that represents a class of replies. The following figure shows the overview of data and levels of
measurement.
Levels of measurement:
The level of measurement specifies the extent to which a number (or word or letter) actually
represents some attribute and, therefore, has implications for the appropriateness of various arithmetic
operations and statistical procedures. There are three levels of measurement—nominal, ordinal, and
interval/ratio—and these levels are paired with qualitative, ranked, and quantitative data, respectively.
Qualitative Data and Nominal Measurement
The first level of measurement is nominal level of measurement. In this level of measurement, the
numbers in the variable are used only to classify the data. In this level of measurement, words, letters,
and alpha-numeric symbols can be used. Suppose there are data about people belonging to different
gender categories. In this case, the person belonging to the female gender could be classified as F and the
person belonging to the male gender could be classified as M. This type of assigning classification is
nominal level of measurement.
The second level of measurement is the ordinal level of measurement. This level of measurement
depicts some ordered relationship among the observations. Suppose a student scores the highest mark of
100 in the class. In this case, he would be assigned the first rank. Then, another classmate scores the
second highest mark of 92; she would be assigned the second rank. A third student scores 81 and he
would be assigned the third rank, and so on. The ordinal level of measurement indicates an ordering of
the measurements.
The third level of measurement is the interval level of measurement. The interval level of
measurement not only classifies and orders the measurements, but it also specifies that the distances
between each interval on the scale are equivalent along the scale from low interval to high interval. A
popular example of this level of measurement is temperature in centigrade, where, for example, the
distance between 940C and 960C is the same as the distance between 1000C and 1020C.
In this level of measurement, the observations, in addition to having equal intervals, can have a
value of zero as well. The zero in the scale makes this type of measurement unlike the other types of
measurement, although the properties are similar to that of the interval level of measurement. Since true
zero is possible, numerical readings reflect the total amount of a person’s weight, and it’s appropriate to
describe one person’s weight as a certain ratio of another’s. It can be said that the weight of a 140 kg
person is twice that of a 70 kg person.
_____________________________________________________________________________________
TYPES OF VARIABLES
Discrete Variables
A discrete variable consists of isolated numbers separated by gaps. Examples include most counts,
such as the number of children in a family, the number of foreign countries you have visited etc.…
Continuous Variables
A continuous variable consists of numbers whose values have no restrictions. Examples include
amounts, such as weights of students, speed of a car etc.…
Approximate Numbers
Whenever values are rounded off, as is always the case with actual values for continuous variables, the
resulting numbers are approximate, never exact. For example, the weights of the male students are
approximate because they have been rounded to the nearest pound. Someone’s weight, in pounds, might
be 140.01438, and so on, to infinity! For practical consideration it should be rounded off.
Independent variable
It’s called “independent” because it’s not influenced by any other variables in the study.
Dependent variable
A dependent variable is the variable that changes as a result of the independent variable manipulation. It
“depends” on independent variable.
This measurement data is used to check whether and to what extent independent variable influences the
dependent variable by conducting statistical analyses.
Confounding variables are a type of peripheral variable that are related to a study’s independent and
dependent variables. A variable must meet two conditions to be a confounder:
1. It must be correlated with the independent variable. This may be a causal relationship, but it does
not have to be.
2. It must be causally related to the dependent variable.
You collect data on sunburns and ice cream consumption. You find that higher ice cream consumption is
associated with a higher probability of sunburn. Does that mean ice cream consumption causes sunburn?
Here, the confounding variable is temperature: hot temperatures cause people to both eat more ice cream
and spend more time outdoors under the sun, resulting in more sunburn.
Observational Studies
An observational study focuses on detecting relationships between variables not manipulated by the
investigator, and it yields less clear-cut conclusions about cause effect relationships than does an
experiment.
_____________________________________________________________________________________
SYLLABUS TOPICS
The table (right) shows one way to organize the performance of students rated in the range of 1 to 10
listed as in table (left). First, arrange a column of consecutive numbers, beginning with the poorest
performance (1) at the bottom and ending with the best performance (10) at the top. Then place a short
vertical stroke or tally next to a number each time its value appears in the original set of data; once this
process has been completed, substitute for each tally count a number indicating the frequency (f) of
occurrence of each value.
A frequency distribution is a collection of observations produced by sorting observations into
classes and showing their frequency (f) of occurrence in each class.
When observations are sorted into classes of single values, the result is referred to as a frequency
distribution for ungrouped data.
Grouped Data:
Guidelines
The “Guidelines for Frequency Distributions” box lists seven rules for producing a well-constructed
frequency distribution. The first three rules are essential and should not be violated. The last four rules are
optional and can be modified or ignored.
Guidelines for Frequency Distributions
Essential
Example: 130–139, 140–149, 150–159, etc. It would be incorrect to use 130–140, 140–150, 150–
160, etc., in which, because the boundaries of classes overlap, an observation of 140 (or 150) could be
assigned to either of two classes.
Example: Listed in Table 2.2 is the class 210–219 and its frequency of zero. It would be incorrect
to skip this class because of its zero frequency.
Example: 130–139, 140–149, 150–159, etc. It would be incorrect to use 130–139, 140–159, etc.,
in which the second class interval (140–159) is twice as wide as the first class interval (130–139).
Optional
4. All classes should have both an upper boundary and a lower boundary.
Example: 240–249. Less preferred would be 240–above, in which no maximum value can be assigned to
observations in this class.
5. Select the class interval from convenient numbers, such as 1, 2, 3 . . . 10, particularly 5 and 10 or
multiples of 5 and 10.
Example: 130–139, 140–149, in which the class interval of 10 is a convenient number. Less
preferred would be 130–142, 143–155, etc., in which the class interval of 13 is not a convenient number.
6. The lower boundary of each class interval should be a multiple of the class interval.
Example: 130–139, 140–149, in which the lower boundaries of 130, 140, are multiples of 10, the
class interval. Less preferred would be 135–144, 145–154, etc., in which the lower boundaries of 135 and
145 are not multiples of 10, the class interval.
Example: The distribution in Table 2.2 uses 12 classes. Less preferred would be the distributions
in Tables 2.3 and 2.4. The distribution in Table 2.3 has too many classes (24), whereas the distribution in
Table 2.4 has too few classes (3).
Gaps between Classes
In well-constructed frequency tables, the gaps between classes, such as between 149 and 150 in Table 2.2,
show clearly that each observation or score has been assigned to one, and only one, class.
The size of the gap should always equal one unit of measurement; that is, it should always equal the
smallest possible difference between scores within a particular set of data. Since the gap is never bigger
than one unit of measurement, no score can fall into the gap
Gaps cannot be ignored when you are determining the actual width of any class interval. The real limits
are located at the midpoint of the gap between adjacent tabled boundaries; that is, one-half of one unit of
measurement below the lower tabled boundary and one-half of one unit of measurement above the upper
tabled boundary.
For example, the real limits for 140–149 in Table 2.2 are 139.5 (140 minus one-half of the unit of
measurement of 1) and 149.5 (149 plus one-half of the unit of measurement of 1), and the actual width of
the class interval would be 10 (from 149.5 - 139.5 = 10).
Constructing Frequency Distributions
Consider the table 1.1 and construct the frequency distributions using the following steps.
1. Find the range, that is, the difference between the largest and smallest observations. The range of
weights in Table 1.1 is 245 133 = 112.
2. Find the class interval required to span the range by dividing the range by the desired number of
classes (ordinarily 10). In the present example
𝑟𝑎𝑛𝑔𝑒 112
𝐶𝑙𝑎𝑠𝑠 𝐼𝑛𝑡𝑒𝑟𝑣𝑎𝑙 = = = 11.2
𝑑𝑒𝑠𝑖𝑟𝑒𝑑 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑐𝑙𝑎𝑠𝑠𝑒𝑠 10
An outlier is a data point that differs significantly from other observations. An outlier may be due
to variability in the measurement or it may indicate experimental error; the latter are sometimes excluded
from the data set. An outlier can cause serious problems in statistical analyses.
Relative frequency distributions show the frequency of each class as a part or fraction of the total
frequency for the entire distribution.
This type of distribution allows us to focus on the relative concentration of observations among different
classes within the same distribution.
To convert a frequency distribution into a relative frequency distribution, divide the frequency for
each class by the total frequency for the entire distribution. Table 2.5 illustrates a relative frequency
distribution based on the weight distribution of Table 2.2.
The conversion to proportions is straightforward. For instance, to obtain the proportion of .06 for
the class 130–139, divide the frequency of 3 for that class by the total frequency of 53. Repeat this
process until a proportion has been calculated for each class.
To convert the relative frequencies in Table 2.5 from proportions to percentages, multiply each
proportion by 100; that is, move the decimal point two places to the right. For example, multiply .06 (the
proportion for the class 130–139) by 100 to obtain 6 percent.
Cumulative frequency distributions show the total number of observations in each class and in
all lower-ranked classes.
This type of distribution can be used effectively with sets of scores, such as test scores for
intellectual or academic aptitude, when relative standing within the distribution assumes primary
importance. Under these circumstances, cumulative frequencies are usually converted, in turn, to
cumulative percentages. Cumulative percentages are often referred to as percentile ranks.
To convert a frequency distribution into a cumulative frequency distribution, add to the frequency
of each class the sum of the frequencies of all classes ranked below it. This gives the cumulative
frequency for that class.
Begin with the lowest-ranked class in the frequency distribution and work upward, finding the
cumulative frequencies in ascending order.
In Table 2.6, the cumulative frequency for the class 130–139 is 3, since there are no classes ranked
lower. The cumulative frequency for the class 140–149 is 4, since 1 is the frequency for that class and 3 is
the frequency of all lower-ranked classes. The cumulative frequency for the class 150–159 is 21, since 17
is the frequency for that class and 4 is the sum of the frequencies of all lower-ranked classes.
Cumulative Percentages
If relative standing within a distribution is particularly important, then cumulative frequencies are
converted to cumulative percentages.
A glance at Table 2.6 reveals that 75 percent of all weights are the same as or lighter than the
weights between 170 and 179 lbs. To obtain this cumulative percentage (75%), the cumulative frequency
of 40 for the class 170–179 should be divided by the total frequency of 53 for the entire distribution.
Percentile Ranks
When used to describe the relative position of any score within its parent distribution, cumulative
percentages are referred to as percentile ranks.
The percentile rank of a score indicates the percentage of scores in the entire distribution with
similar or smaller values than that score. Thus a weight has a percentile rank of 80 if equal or lighter
weights constitute 80 percent of the entire distribution.
The percentile formula determines the performance of a person over others. The percentile
formula is used in finding where a student stands in the test compared to other candidates. A percentile is
a number where a certain percentage of scores fall below the given number.
The percentile of x is the ratio of the number of values below x to the total number of values
multiplied by 100. i.e., the percentile formula is
The assignment of exact percentile ranks requires that cumulative percentages be obtained from
frequency distributions for ungrouped data. If we have access only to a frequency distribution for grouped
data, as in Table 2.6, cumulative percentages can be used to assign approximate percentile ranks. In Table
2.6, for example, any weight in the class 170–179 could be assigned an approximate percentile rank of 75,
since 75 is the cumulative percent for this class.
When, among a set of observations, any single observation is a word, letter, or numerical code, the data
are qualitative. Frequency distributions for qualitative data are easy to construct. Simply determine the
frequency with which observations occupy each class, and report these frequencies as shown in Table 2.7
for the Facebook profile survey (Whether person is possessing Facebook account or not).
Ordered Qualitative Data
Qualitative data have an ordinal level of measurement because observations can be ordered from
least to most, that order should be preserved in the frequency table, as illustrated in Table 2.8, in which
military ranks are listed in descending order from general to lieutenant.
Frequency distributions for qualitative variables can always be converted into relative frequency
distributions, as illustrated in Table 2.8. Furthermore, if measurement is ordinal because observations can
be ordered from least to most, cumulative frequencies (and cumulative percentages) can be used.
For example, that a captain has an approximate percentile rank of 63 among officers since 62.5 (or
63) is the cumulative percent for this class. If measurement is only nominal because observations cannot
be ordered, as in Table 2.7, a cumulative frequency distribution is meaningless.
Graphs
Data can be described clearly and concisely with the aid of a well-constructed frequency
distribution. And data can often be described even more vividly, particularly to communicate with a
general audience, by converting frequency distributions into graphs.
Graphs for Quantitative Data
Histograms
Histogram is a bar-type graph for quantitative data. The common boundaries between adjacent
bars emphasize the continuity of the data, as with continuous variables.
The following histogram shows the weight distribution of male students. A casual glance at this
histogram confirms previous conclusions: a dense concentration of weights among the 150s, 160s, and
170s, with a spread in the direction of the heavier weights.
• Equal units along the horizontal axis (the X axis, or abscissa) reflect the various class intervals of
the frequency distribution.
• Equal units along the vertical axis (the Y axis, or ordinate) reflect increases in frequency. (The
units along the vertical axis do not have to be the same width as those along the horizontal axis.)
• The intersection of the two axes defines the origin at which both numerical scales equal 0.
• Numerical scales always increase from left to right along the horizontal axis and from bottom to
top along the vertical axis.
• The body of the histogram consists of a series of bars whose heights reflect the frequencies for
the various classes.
Frequency Polygon
B. Place dots at the midpoints of each bar top or, in the absence of bar tops, at midpoints for
classes on the horizontal axis, and connect them with straight lines.[To find the midpoint of any
class, such as 160–169, simply add the two tabled boundaries (160 + 169 = 329) and divide this
sum by 2 (329/2 = 164.5).]
C. Anchor the frequency polygon to the horizontal axis. First, extend the upper tail to the midpoint
of the first unoccupied class (250–259) on the upper flank of the histogram. Then extend the lower
tail to the midpoint of the first unoccupied class (120–129) on the lower flank of the histogram.
Now all of the area under the frequency polygon is enclosed completely.
D. Finally, erase all of the histogram bars, leaving only the frequency polygon.
Frequency polygons are particularly useful when two or more frequency distributions or relative
frequency distributions are to be included in the same graph.
Another technique for summarizing quantitative data is a stem and leaf display. It is a device for
sorting quantitative data on the basis of leading and trailing digits. Stem and leaf displays are ideal for
summarizing distributions without destroying the identities of individual observations.
Constructing a Display
The leftmost panel of Table 2.9 re-creates the weights of the 53 male statistics students listed in
Table 1.1.
• To construct the stem and leaf display for these data, first note the weights range. i.e. from
the 130s to the 240s.
• Arrange a column of numbers, the stems, beginning with 13 (representing the 130s) and
ending with 24 (representing the 240s).
• Draw a vertical line to separate the stems, which represent multiples of 10, from the space
to be occupied by the leaves, which represent multiples of 1.
• Next, enter each raw score into the stem and leaf display. As suggested by the shaded
coding in Table 2.9, the first raw score of 160 reappears as a leaf of 0 on a stem of 16.
• The next raw score of 193 reappears as a leaf of 3 on a stem of 19, and the third raw score
of 226 reappears as a leaf of 6 on a stem of 22, and so on, until each raw score reappears as
a leaf on its appropriate stem.
The weight data have been sorted by the stems. All weights in the 130s are listed together; all of
those in the 140s are listed together, and so on. This simple method only works if, as in the present
display, stem values are listed from smallest at the top to largest at the bottom.
Selection of Stems
Stem values are not limited to units of 10. Depending on the data stem value can be varied as 1,
100, 1000, or even .1, .01, .001, and so on. For instance, an annual income of $23,784 could be displayed
as a stem of 23 (thousands) and a leaf of 784. (Leaves consisting of two or more digits, such as 784, are
separated by commas.)
Typical Shapes
An important characteristic of a frequency distribution is its shape. Figure 2.3 shows some of the
more typical shapes for smoothed frequency polygons.
Normal Distribution
A normal distribution, sometimes called the bell curve is a distribution that occurs naturally in
many situations. For example, the bell curve is seen in tests like the SAT and GRE. The bulk of students
will score the average (C), while smaller numbers of students will score a B or D. An even smaller
percentage of students score an F or an A. This creates a distribution that resembles a bell.
Bimodal
Data distributions in statistics can have one peak, or they can have several peaks. The type of
distribution you might be familiar with seeing is the normal distribution, or bell curve, which has one
peak. The bimodal distribution has two peaks.
Example: Scores fall into bimodal distribution with one group of students getting scores between 70 to
75 marks out of 100 and another group of students getting scores between 25 to 30 marks.
• A lopsided distribution that includes a few extreme observations in the positive direction (to the
right of the majority of observations) is positively skewed.
• A lopsided distribution caused by a few extreme observations in the negative direction (to the left
of the majority of observations) is negatively skewed.
A glance at this graph confirms that Yes replies occur approximately twice as often as No replies.
As with histograms, equal segments along the horizontal axis are allocated to the different words
or classes that appear in the frequency distribution for qualitative data. Likewise, equal segments along
the vertical axis reflect increases in frequency. The body of the bar graph consists of a series of bars
whose heights reflect the frequencies for the various words or classes.
Gaps are placed between adjacent bars of bar graphs to emphasize the discontinuous nature of
qualitative data.
Constructing Graphs
___________________________________________________________________________________
Describing Data with Averages Averages consist of numbers (or words) about which the data are, in
some sense, centred. They are often referred to as measures of central tendency, the several types of
average yield numbers or words that attempt to describe, most generally, the middle or typical value for a
distribution.
• Mode
• Median
• Mean
Mode
The mode reflects the value of the most frequently occurring score.
A mode is defined as the value that has a higher frequency in a given set of values. It is the value that
appears the most number of times.
Example: In the given set of data: 2, 4, 5, 5, 6, 7, the mode of the data set is 5 since it has appeared in the
set twice.
Distributions can have more than one mode (or no mode at all). Distributions with two obvious peaks,
even though they are not exactly the same height, are referred to as bimodal.
Median
The median reflects the middle value when observations are ordered from least to most.
The median splits a set of ordered observations into two equal parts, the upper and lower halves.
In other words, the median has a percentile rank of 50, since observations with equal or smaller values
constitute 50 percent of the entire distribution.
Table 3.2 shows how to find the median for two different sets of scores. The following steps are followed
to find the median.
The median term describes the middle-ranked term; the modal term describes the most
frequent term in the distribution.
Mean
The mean is the most common average, one you have doubtless calculated many times. The mean is
found by adding all scores and then dividing by the number of scores.
That is,
Example:
In a class there are 20 students and they have secured a percentage of 88, 82, 88, 85, 84, 80, 81, 82, 83,
85, 84, 74, 75, 76, 89, 90, 89, 80, 82, and 83. Find the mean percentage obtained by the class.
= [88 + 82 + 88 + 85 + 84 + 80 + 81 + 82 + 83 + 85 + 84 + 74 + 75 + 76 + 89 + 90
+ 89 + 80 + 82 + 83]/20
= 1660/20
= 83
Hence, the mean percentage of each student in the class is 83%.
Statisticians distinguish between two types of means—the population mean and the sample
mean—depending on whether the data are viewed as a population (a complete set of scores) or as a
sample (a subset of scores).
It’s usually more efficient to substitute symbols for words in statistical formulas, including the
word formula given above for the mean. When symbols are used, 𝑋̅ designates the sample mean, and the
formula becomes
∑𝑋
𝑋̅ =
𝑛
The formula for the population mean differs from that for the sample mean only because of a
change in some symbols. In statistics, Greek symbols usually describe population characteristics, such as
the population mean, while English letters usually describe sample characteristics, such as the sample
mean. The population mean is represented by μ (pronounced “mu”), the lowercase Greek letter m for
mean,
∑𝑋
𝜇=
𝑁
where the uppercase letter N refers to the population size. Otherwise, the calculations are the same as
those for the sample mean.
The mean reflects the values of all scores, not just those that are middle ranked (as with the median),
or those that occur most frequently (as with the mode).
Mode Always Appropriate for Qualitative Data For quantitative data all the three averages can be
used. But when the data are qualitative, the choice among averages is restricted. The mode always can be
used with qualitative data. For instance, Yes qualifies as the modal or most typical response for the
Facebook profile question (Table 2.7).
Median Sometimes Appropriate
The median can be used whenever it is possible to order qualitative data from least to most
because the level of measurement is ordinal.
It’s easiest to determine the median class for ordered qualitative data by using relative frequencies, as in
Table 3.5.
Cumulate the relative frequencies, working up from the bottom of the distribution, until the
cumulative percentage first equals or exceeds 50 percent. Since the corresponding class includes the
median and, roughly speaking, splits the distribution into an upper and a lower half, it is designated as the
median or middle-ranked class.
For example, the qualitative data in Table 3.5 can be ordered from lieutenant to general. Starting
at the bottom of Table 3.5 and cumulating upward, 25. 5 percent is obtained for the class of lieutenant and
a cumulative percent 62.5 for the class of captain. Accordingly, since it includes a cumulative percent of
50, captain is the median rank of officers in the U.S. Army.
Inappropriate Averages
It would not be appropriate to report a median for unordered qualitative data with nominal
measurement. It is inappropriate to report a mean for any qualitative data, such as the ranks of officers in
the U.S. Army. After all, words cannot be added and then divided, as required by the formula for the
mean.
Describing Variability
Measures of variability, that is, measures of the amount by which scores are dispersed or scattered
in a distribution.
1. Standard deviation
2. Variance
3. Range
4. Interquartile range
In Figure 4.1, each of the three frequency distributions consists of seven scores with the same
mean (10) but with different variabilities like distribution A has the least variability, distribution B has
intermediate variability, and distribution C has the most variability.
Range
One of the measures of variability and essential tools in statistics is the range. The range is the
difference between the largest and smallest scores.
In Figure 4.1, distribution A, the least variable, has the smallest range of 0 (from 10 to 10);
distribution B, the moderately variable, has an intermediate range of 2 (from 11 to 9); and distribution C,
the most variable, has the largest range of 6 (from 13 to 7), in agreement with our intuitive judgments
about differences in variability.
The range is a handy measure of variability that can readily be calculated and understood.
Shortcomings
1. First, since its value depends on only two Scores —the largest and the smallest—it fails to use the
information provided by the remaining scores.
2. The value of the range tends to increase with increases in the total number of scores.
Variance
The variance and particularly its square root, the standard deviation, serve as key components for
other important statistical measures.
The term variance refers to a statistical measurement of the spread between numbers in a data set.
More specifically, variance measures how far each number in the set is from the mean (average), and thus
from every other number in the set.
In the case of the variance, each original score is re-expressed as a distance or deviation from the
mean by subtracting the mean.
For each of the three distributions in Figure 4.1, the face values of the seven original scores
(shown as numbers along the X axis) have been re-expressed as deviation scores from their mean of 10
(shown as numbers in the boxes). For example, in distribution C, one score coincides with the mean of
10, four scores (two 9s and two 11s) deviate 1 unit from the mean, and two scores (one 7 and one 13)
deviate 3 units from the mean, yielding a set of seven deviation scores: one 0, two –1s, two 1s, one –3,
and one 3. (Deviation scores above the mean are assigned positive signs; those below the mean are
assigned negative signs.)
Standard Deviation
The square root of the mean of all squared deviations from the mean, that is,
The sum of squares, symbolized by SS can be defined by two formulas the definition formula,
which is easier to understand and remember, and the computation formula, which usually is more
efficient.
The definition formula provides the most accessible version of the population sum of squares:
where SS represents the sum of squares, Σ directs us to sum over the expression to its right, and
(X − μ)2 denotes each of the squared deviation scores.
1. Subtract the population mean, μ, from each original score, X, to obtain a deviation score, X − μ.
When the mean equals some complex number, such as 169.51, or the number of scores is large the
more efficient computation formula can be used:
where Σ X2 , the sum of the squared X scores, is obtained by first squaring each X score and then
summing all squared X scores; (Σ X)2 , the square of sum of all X scores, is obtained by first adding all X
scores and then squaring the sum of all X scores; and N is the population size.
𝑠𝑠
8. Take the square root of σ 2 to obtain the population standard deviation, 𝜎 = √ 𝑁 .
Sample notation can be substituted for population notation in the above two formulas without causing any
essential changes:
or, in symbols:
where the squared lowercase Greek letter, σ2 (pronounced “sigma squared”), represents the population
variance, SS is the sum of squared deviations for the population, and N is the population size.
The standard deviation is obtained by taking the square root of the variance, that is,
where σ represents the population standard deviation, instructs us to take the square root of the covered
expression.
Although the sum of squares term remains essentially the same for both populations and samples,
there is a small but important change in the formulas for the variance and standard deviation for samples.
This change appears in the denominator of each formula where N, the population size, is replaced by the
sample size, n − 1, as shown:
where s2 and s represent the sample variance and sample standard deviation, SS is the sample sum
of squares.
Calculation of Sample Standard Deviation (S) (Definition Formula)
Degrees of freedom is a very important notion in inferential statistics. Degrees of freedom (df)
refers to the number of values that are free to vary, given one or more mathematical restrictions, in a
sample being used to estimate a population characteristic.
The concept of degrees of freedom is introduced only because scores in a sample are used to
estimate some unknown characteristic of the population. Typically, when used as an estimate, not all
observed values in the sample are free to vary because of one or more mathematical restrictions.
As has been noted, when n deviations about the sample mean are used to estimate variability in
the population, only n − 1 are free to vary. As a result, there are only n − 1 degrees of freedom, that is, df
= n − 1.
We can use degrees of freedom to rewrite the formulas for the sample variance and standard
deviation:
Interquartile Range (IQR)
The most important spinoff of the range, the interquartile range (IQR), is simply the range for the
middle 50 percent of the scores. More specifically, the IQR equals the distance between the third quartile
(or 75th percentile) and the first quartile (or 25th percentile), that is, after the highest quarter (or top 25
percent) and the lowest quarter (or bottom 25 percent) have been trimmed from the original set of scores.
Qualitative Data
Measures of variability are virtually non-existent for qualitative or nominal data. It is probably
adequate to note merely whether scores are evenly divided among the various classes (maximum
variability), unevenly divided among the various classes (intermediate variability), or concentrated mostly
in one class (minimum variability).
If qualitative data can be ordered because measurement is ordinal (or if the data are ranked), then
it’s appropriate to describe variability by identifying extreme scores (or ranks).
-------------------------------------------------------------------------------------------------------------------------------
More accurate generalizations usually can be obtained from distributions based on larger numbers
of data points.
For example: A distribution based on 30,910 men usually is more accurate than one based on 3091, and a
distribution based on 3,091,000 usually is even more accurate.
Normal curve has been superimposed on the original distribution. Any generalizations based on
the smooth normal curve will tend to be more accurate than those based on the original distribution
In Figure 5.2, the idealized normal curve has been superimposed on the original distribution for
3091 men. Irregularities in the original distribution, most likely due to chance, are ignored by the smooth
normal curve.
Accordingly, any generalizations based on the smooth normal curve will tend to be more accurate
than those based on the original distribution.
The total area under the normal curve in Figure 5.2 can be identified with all data points (Eg.
heights of all men applicants for FBI jobs). Viewed relative to the total area, the shaded area represents
the proportion of applicants who will be eligible because they are shorter than exactly 66 inches.
1. Obtained from a mathematical equation, the normal curve is a theoretical curve defined for a
continuous variable, and is symmetrical bell-shaped curve.
2. Because the normal curve is symmetrical, its lower half is the mirror image of its upper half.
3. Being bell shaped, the normal curve peaks above a point midway along the horizontal spread and
then tapers off gradually in either direction from the peak
4. The values of the mean, median (or 50th percentile), and mode, located at a point midway along
the horizontal spread, are the same for the normal curve.
The various types of normal curves can be produced by an arbitrary change in the value of either
the mean (μ) or the standard deviation (σ).
For example, changing the mean height from 69 to 79 inches produces a new normal curve that, as shown
in panel A of Figure 5.3, is displaced 10 inches to the right of the original curve. Dramatically new
normal curves are produced by changing the value of the standard deviation. As shown in panel B of
Figure 5.3, changing the standard deviation from 3 to 1.5 inches produces a more peaked normal curve
with smaller variability, whereas changing the standard deviation from 3 to 6 inches produces a shallower
normal curve with greater variability.
The differences in appearance among normal curves are less important. Because of their common
mathematical origin, every normal curve can be interpreted in exactly the same way once any distance
from the mean is expressed in standard deviation units.
Once any distance from the mean has been expressed in standard deviation units, the standard normal
table can be used to determine the corresponding proportion of the area under the normal curve.
Z Scores
A z score is a unit-free, standardized score that, regardless of the original units of measurement,
indicates how many standard deviations a score is above or below the mean of its distribution.
To obtain a z score, express any original score, whether measured in inches, milliseconds, dollars,
IQ points, etc., as a deviation from its mean (by subtracting its mean) and then split this deviation into
standard deviation units (by dividing by its standard deviation), that is,
𝑋−𝜇
𝑧=
𝜎
where X is the original score and μ and σ are the mean and the standard deviation, respectively,
for the normal distribution of the original scores. Since identical units of measurement appear in both the
numerator and denominator of the ratio for z, the original units of measurement cancel each other and the
z score emerges as a unit-free or standardized number, often referred to as a standard score.
2. A number indicating the size of its deviation from the mean in standard deviation units.
A z score of 2.00 always signifies that the original score is exactly two standard deviations above
its mean. Similarly, a z score of –1.27 signifies that the original score is exactly 1.27 standard deviations
below its mean. A z score of 0 signifies that the original score coincides with the mean.
When X with 66 (the maximum permissible height), μ with 69 (the mean height), and σ with 3
(the standard deviation of heights) and solve for z:
𝑋−𝜇 66−69 −3
The formula for Z score 𝑧 = = = = -1
𝜎 3 3
This shows the score is exactly one standard deviation below the mean.
The standard normal curve is a special normal distribution where the mean is 0 and the standard
deviation is 1.
To verify (rather than prove) that the mean of a standard normal distribution equals 0, replace X in
the z score formula with μ, the mean of any (nonstandard) normal distribution, and then solve for z:
𝑋−𝜇 𝜇−𝜇 0
𝑀𝑒𝑎𝑛 𝑜𝑓 𝑧 = = = =0
𝜎 𝜎 𝜎
Likewise, to verify that the standard deviation of the standard normal distribution equals 1, replace
X in the z score formula with μ + 1σ, the value corresponding to one standard deviation above the mean
for any (nonstandard) normal distribution, and then solve for z:
𝑋 − 𝜇 𝜇 + 1𝜎 − 𝜇 1𝜎
𝑆𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑜𝑓 𝑧 = = = =1
𝜎 𝜎 𝜎
Although there are an infinite number of different normal curves, each with its own mean and standard
deviation, there is only one standard normal curve, with a mean of 0 and a standard deviation of 1.
Figure 5.4 illustrates the emergence of the standard normal curve from three different normal
curves: that for the men’s heights, with a mean of 69 inches and a standard deviation of 3 inches; that for
the useful lives of 100-watt electric light bulbs, with a mean of 1200 hours and a standard deviation of
120 hours; and that for the IQ scores of fourth graders, with a mean of 105 points and a standard deviation
of 15 points.
To find the proportion for the shaded areas in Figure 5.4, the same z score of –1.00 can be used
when referring to the table for the standard normal curve, the one table for all normal curves.
Standard Normal Table
Essentially, the standard normal table consists of columns of z scores coordinated with columns of
proportions.
▪ Columns are arranged in sets of three, designated as A, B, and C in the legend at the top of the
table.
▪ When using the top legend, all entries refer to the upper half of the standard normal curve.
▪ The entries in column A are z scores.
▪ Columns B and C indicate how the z score splits the area in the upper half of the normal curve.
▪ Column B indicates the proportion of area between the mean and the z score.
▪ Column C indicates the proportion of area beyond the z score, in the upper tail of the standard
normal curve.
Columns are designated as A′, B′, and C′ in the legend at the bottom of the table.
All entries refer to the lower half of the standard normal curve.
Column B′ indicates the proportion of area between the mean and the negative z score.
Column C′ indicates the proportion of area beyond the negative z score, in the lower tail of the
standard normal curve.
There are two main types of normal curve problems. In the first type of problem, a known score
(or scores) is used to find an unknown proportion. In the second type of problem, the procedure is
reversed i.e. a known proportion is used to find an unknown score (or scores).
When using the standard normal table, it is important to remember that for any z score, the
corresponding proportions in columns B and C (or columns B′ and C′) always sum to .5000. Similarly,
the total area under the normal curve always equals 1.0000, the sum of the proportions in the lower and
upper halves, that is, .5000 + .5000. Figure 5.5 summarizes how to interpret the normal curve table.
Finding Proportions
Find the proportion of all FBI applicants who are shorter than exactly 66 inches, given that the
distribution of heights approximates a normal curve with a mean of 69 inches and a standard deviation of
3 inches.
The answer will be obtained from column C′ of the standard normal table, since the target
area coincides with the type of area identified with column C′, that is, the area in the lower tail beyond
a negative z.
It can be concluded that only .1587 (or .16) of all of the FBI applicants will be shorter than 66
inches.
Assume that high school students’ IQ scores approximate a normal distribution with a mean of 105 and
a standard deviation of 15. What proportion of IQs are more than 30 points either above or below the
mean?
1. Sketch a normal curve and shade in the two target areas, as in the top panel of Figure 5.8.
2. Plan your solution according to the normal table.
The solution to this type of problem is straightforward because each of the target areas can be read
directly from standard normal table. The target area in the tail to the right can be obtained from
column C, and that in the tail to the left can be obtained from column C′.
3. Convert X to z by expressing IQ scores of 135 and 75 as
75 − 105 −30
𝑧= = = −2.00
15 15
135 − 105 30
𝑧= = = 2.00
15 15
4. Find the target area.
In standard normal table, locate a z score of 2.00 in column A, and note the corresponding
proportion of .0228 in column C. Because of the symmetry of the normal curve, there is no need to
enter the table again to find the proportion below a z score of –2.00. Instead, merely double the
above proportion of .0228 to obtain .0456, which represents the proportion of students with IQs
more than 30 points either above or below the mean.
Finding Scores
Exam scores for a large psychology class approximate a normal curve with a mean of 230 and a standard
deviation of 50. Furthermore, students are graded “on a curve,” with only the upper 20 percent being
awarded grades of A. What is the lowest score on the exam that receives an A?
1. Sketch a normal curve and, on the correct side of the mean, draw a line representing the target score,
as in Figure 5.9.
2. Plan your solution according to the normal table.
Since the target score is on the right side of the mean, concentrate on the area in the upper half of
the normal curve, as described in columns B and C.
The following formula is used in the conversion of z score to the target score.
Assume that the annual rainfall in the San Francisco area approximates a normal curve with a mean of
22 inches and a standard deviation of 4 inches. What are the rainfalls for the more atypical years,
defined as the driest 2.5 percent of all years and the wettest 2.5 percent of all years?
1. Sketch a normal curve.
On either side of the mean, draw two lines representing the two target scores, as in Figure 5.10.
The smaller (driest) target score splits the total area into .0250 to the left and .9750 to the right, and
the larger (wettest) target score does the exact opposite.
Important Note: For additional problems refer page no. 103 - 106 in II Text Book.