THE NORMAL DISTRIBUTION
The normal distribution is central to the study of Statistics, as it serves as the basis for solving various
types of statistical problems. In most cases, the distributions of variables such as students' grades,
people's weights or heights, families' incomes, or children's IQs are approximately normal. If we take the
heights of adults as an example, the normal distribution means, roughly, that:
• There are relatively few short adults
• There are relatively few tall adults
• The height of most adults will tend towards a middle value (the average value) between the
shortest and the tallest.
In the histogram shown in Figure 1, the height of each rectangle represents the frequency, or, in our
example, the number of adults. Our histogram then tells the story at a glance – few short adults, few tall
adults, and many adults with average heights.
many
few few
short average tall
Figure 1. Histogram of heights of adults in the example.
Skewed Distributions
While many variables tend to follow a normal distribution, some do not. We call such types skewed
distributions. Skewed distributions are of two types: (1) skewed to the right, and (2) skewed to the left.
A distribution that is skewed to the right has a longer right tail, while one that is skewed to the left has a
longer left tail. The distribution in Figure 2 has no skewness because it is symmetric. Figure 3 shows a
distribution that is skewed left, or negatively skewed, and Figure 4 shows a distribution that is skewed
right, or positively skewed.
Figure 2 Figure 3 Figure 4
Vicente G. Padilla
Associate Professor IV, CatSU College of Science
Properties of a Normal Curve
The normal distribution exhibits the following characteristics:
1. It is a continuous distribution.
2. It is a symmetrical distribution about its mean.
3. It is asymptotic to the horizontal axis.
4. It is unimodal.
5. It is a family of curves.
6. Area under the curve is 1.
The normal distribution is symmetrical. Each half of the distribution is a mirror image of the
other half. Many normal distribution tables contain probability values for only one side of the distribution
are identical because of symmetry. In theory, the normal distribution approaches the horizontal axis
asymptotically. That is, it does not touch the x-axis and extends infinitely in both directions. The reality
is that most applications of the normal curve involve experiments with finite sets of potential outcomes.
The normal curve, also called the bell-shaped curve, is sometimes referred to as the bell curve. It
is unimodal, with values clustering in a single region of the graph, centered on the curve. The normal
distribution actually is a family of curves. Every unique value of the mean and every unique value of the
standard deviation result in a different normal curve. In addition, the total area under any normal
distribution is 1. The area under the curve yields the probabilities, so the total probability for a normal
distribution is 1. Because the area is symmetric, the area of the distribution on each side of the mean is
0.5.
The Probability Density Function of the Normal Distribution
The normal distribution is characterized by two parameters: the mean,, and the standard deviation,.
The values of and produce a normal distribution. The density function of the normal distribution is
1
𝑒 −1/2[(𝑥−)/)]
2
𝑓(𝑥) =
√2
where
= mean of x
= standard deviation of x
= 3.14159…, and
e = 2.71828…
Standardized Normal Distribution
In making use of the properties of the normal curve to solve certain types of statistical problems, one must
first learn how to find areas under the normal curve. The first step in finding areas under the normal
curve is to convert the normal curve of any given variable into a standardized normal curve by using the
formula:
𝑥 −
𝑧=
where
z = standard score
= mean
= standard deviation
x = a given value of a particular variable
Vicente G. Padilla
Associate Professor IV, CatSU College of Science
A z-score is the number of standard deviations that a value x is above or below the mean. If the
value of x is less than the mean, the z score is negative; if the value of x is more than the mean, the z score
is positive; and if the value of x equals the mean, the associated z score is zero. This formula converts the
distance of any x value from its mean into standard deviation units. A standard z-score table can be used
to find probabilities for any normal curve problem that has been converted to z scores.
The z distribution is a normal distribution with a mean of 0 and a standard deviation of 1. Any value of
x at the mean of a normal curve is zero standard deviations from the mean. Any value of x that is one
standard deviation above the mean has a z value of 1. The graph of the normal distribution's density
function is shown in Figure 5. In this graph, we have indicated the areas within 1, 2, and 3 standard
deviations of the mean (i.e., between z = -1 and +1, z = -2 and +2, z = -3 and +3) as equal, respectively, to
68.27%, 95.45% and 99.73% of the total area, which is one. This means that
𝑃(−1 𝑍 1) = 0.6827, 𝑃(−2 𝑍 2) = 0.9545, 𝑃(−3 𝑍 3) = 0.9973,
Figure 5
Table 1 gives the total area under the z curve between 0 and any point on the positive z-axis.
Since the curve is symmetric, the area under the curve between z and 0 is the same whether z is positive or
negative (the sign on the z value designates whether the z score is above or below the mean). The table
areas or probabilities are always positive.
Solving Normal Curve Problems
The mean and standard deviation of a normal distribution, and the z formula and table, enable a
researcher to determine the probabilities for intervals of any particular values of a normal curve. One
example is the many possible probability values of GMAT scores.
Figure 6
The Graduate Management Admission Test (GMAT),
produced by the Princeton Review in Princeton, New Jersey, is widely
used by U.S. business schools as an admissions requirement.
Assuming the scores are normally distributed, the probability of
achieving scores within various GMAT score ranges can be
determined. In a recent year, the mean GMAT score was 540, and the
standard deviation was about 100. What is the probability that a
randomly selected score from this distribution of the GMAT is
between 660 and the mean? That is,
𝑃(540 ≤ 𝑥 ≤ 660 = 540 and = 100) =?
Figure 6 is a graphical representation of this problem.
Vicente G. Padilla
Associate Professor IV, CatSU College of Science
Table 1
Normal Distribution
The z formula yields the number of standard deviations by which the x value, 660, is from the
mean.
𝑥− 660 −540 120
𝑧= = = = 1.20
100 100
The z value of 1.20 indicates that the GMAT score of 660 is 1.2 standard deviations above the
mean. The z distribution values in Table 1 give the probability that a value lies between x and the mean.
The whole-number and tenths-place portion of the z score appears in the first column of Table 1 (the 1.0
portion of this z score). Across the top of the table are the values of the hundredths-place portion of the z
score. For this score, the hundredths-place value is 0. The probability value in Table 1 for z = 1.20 is
.3849. The shaded portion of the curve at the top of the table indicates that the probability value reported
is always the probability or area to the left of the mean. In this particular example, that is the desired area.
Vicente G. Padilla
Associate Professor IV, CatSU College of Science
Thus, 0.3849 of the GMAT scores are between 660 and the mean of 540. Figure 7(a) depicts the solution
graphically in terms of x values. Figure 7(b) shows the solution in terms of z values.
Figure 7
Graphical solution
to the GMAT
problem
Demonstration Problem 1. What is the probability of obtaining a score greater than 750 on a GMAT
test that has a mean of 540 and a standard deviation of 100? Assume GMAT scores are normally
distributed.
𝑃(𝑥 > 750 = 540 and = 100) = ?
Solution
Examine the diagram at the right.
This problem calls for determining the area of the upper tail
of the distribution. The z score for this problem is
𝑥− 750−540 210
𝑧= = = = 2.10
100 100
Table 1 gives the probability of .4821 for this z score. This value is the probability of randomly
drawing a GMAT with a score between the mean and 750. Finding the probability of getting a score
greater than 750, which is the tail of the distribution, requires subtracting the probability value of .4821
from .5000, because each half of the distribution contains .5000 of the area. The result is .0179. Note
that an attempt to determine the area of x 750 instead of x > 750 would have made no difference
because, in continuous distributions, the area under an exact number such as x = 750 is zero. A line
segment has no width and hence no area.
.5000 (probability of x greater than the mean)
−.4821 (probability of x between 750 and the mean)
.0179 (probability of x greater than 750)
The solution is depicted graphically in (a) for x values and (b) for z values.
Vicente G. Padilla
Associate Professor IV, CatSU College of Science
Demonstration Problem 2. For the same GMAT examination, what is the probability of randomly
drawing a score that is 590 or less?
𝑃(𝑥 590 = 540 and = 100) = ?
Solution
A sketch of this problem is shown at the right. Determine the
area under the curve for all values less than or equal to 590.
The z value for the area between 590 and the mean is equal to
𝑥− 590−540 50
𝑧= = = = 0.50
100 100
The area under the curve for z = 0.50 is .1915, which
is the probability of getting a score between 590 and the mean. However, obtaining the probability for all
values less than or equal to 590 also requires including the values less than the mean. Because one-half or
.5000 of the values are less than the mean, the probability of x 590 is found as follows.
.5000 (probability of values less than the mean)
+.1915 (probability of values between 590 and the mean)
.6915 (probability of values 590)
This solution is depicted graphically in (a) for x values and in (b) for z values.
Demonstration Problem 3. What is the probability of randomly obtaining a score between 350 and 630
on the GMAT exam?
𝑃(350 < 𝑥 < 630 = 540 and = 100) = ?
Solution
The following sketch depicts the problem graphically at the
right: determine the area between x = 350 and x = 630,
which spans the mean value. Because areas in the z
distribution are given in relation to the mean, this problem
must be worked as two separate problems and the results
combined.
A z score is determined for each x value.
𝑥− 630−540 90
𝑧= = = = 0.90
100 100
and
𝑥− 350−540 −190
𝑧= = = = −1.90
100 100
Vicente G. Padilla
Associate Professor IV, CatSU College of Science
Note that this z value (z = -1.90) is negative. A negative z value indicates that the x value is below the
mean and the z value is on the left side of the distribution. None of the z values in Table 1 is negative.
However, because the normal distribution is symmetric, probabilities for z values on the left side of the
distribution are the same as the values on the right side of the distribution. The negative sign in the z
value merely indicates that the area is on the left side of the distribution. The probability is always
positive.
The probability for z = 0.90 is .3159; the probability for z = -1.90 is .4713. the solution of
𝑃(350 < 𝑥 < 630) is obtained by summing the probabilities.
.3159 (probability of a value between the mean and 630)
+.4713 (probability of a value between the mean and 350)
.7872 (probability of a value between 350 and 630)
Graphically, the solution is shown in (a) for x values and in (b) for z values.
Demonstration Problem 4. What is the probability of getting a score between 400 and 500 on the same
GMAT exam?
𝑃(400 < 𝑥 < 500 = 540 and = 100) = ?
Solution
The following sketch reveals that the solution to the problem
involves determining the area of the shaded slice in the lower
half of the curve.
In this problem, the two x values are on the same side
of the mean. The probability for each x value must be
determined, and the final probability found by subtracting the
two areas.
𝑥− 400−540 −140
𝑧= = = = −1.40
100 100
and
𝑥− 500−540 −40
𝑧= = = = −0.40
100 100
The probability associated with z = -1.40 is .4192.
The probability associated with z = -0.40 is .1554.
Subtracting gives the solution.
.4192 (probability of a value between 400 and the mean)
-.1554 (probability of a value between 500 and the mean)
.2638 (probability of a value between 400 and 500)
Vicente G. Padilla
Associate Professor IV, CatSU College of Science
Graphically, the solution is shown in (a) for x values and in (b) for z values.
Demonstration Problem 5. Runzheimer International publishes business travel cost data for cities
worldwide. In particular, they publish per diem totals, which represent the average costs for the typical
business traveler, including three meals a day in business-class restaurants and single-rate lodging in
business-class hotels and motels. If 86.65% of the per diem costs in Buenos Aires, Argentina, are less
than $449 and if the standard deviation of per diem costs is $36, what is the average per diem cost in
Buenos Aires? Assume that per diem costs are normally distributed.
Solution
In this problem, the standard deviation and an x-value are given; the objective is to determine the value of
the mean. Examination of the z-score formula reveals four variables: x, , , and z. In this problem, only
two of the four variables are given. Because solving one equation with two unknowns is impossible, one
of the other unknowns must be determined. The value of z can be determined from the standard
distribution table (Table 1).
Because 86.65% of the values are less than x = $449,
36.65% of the per diem costs are between $449 and the mean.
The other 50% of the per diem costs are in the lower half of
the distribution. Converting the percentage to a proportion
yields .3665 of the values between the x value and the mean.
What z value is associated with this area? This area, or
probability, of .3665 in Table 1 is associated with the z value
of 1.11. This z-value is positive because it lies in the upper
half of the distribution.
Using the z value of 1.11, the x value of $449, and the value of $36, the mean can be solved for
algebraically.
𝑥 −
𝑧=
$449 −
1.11 =
$36
and
= $499 – ($36)(1.11) = $449 - $39.96 = $409.04
The mean per diem cost for business travel in Buenos Aires is $409.04.
Vicente G. Padilla
Associate Professor IV, CatSU College of Science
Demonstration Problem 6. The U.S. Environmental Protection Agency publishes figures on solid waste
generation in the United States. One year, the average number of waste generated per person per day was
3.58 pounds. Suppose the daily amount of waste generated per person is normally distributed, with a
standard deviation of 1.04 pounds. Of the daily amounts of waste generated per person, 67.72% would be
greater than what amount?
Solution
The mean and standard deviation are given, but x and z are unknown. The problem is to find a specific x
value such that .6772 of the x values are greater than that value.
If .6772 of the values are greater than x, then .1772 are between x and the mean (.6772 - .5000).
Table 1 shows that the probability of .1772 is associated with a z value of 0.46. Because x is less than the
mean, the z value actually is -0.46. Whenever an x value is less than the mean, its associated z value is
negative and should be reported that way.
Solving the z equation yields
𝑥 −
𝑧=
𝑥 − 3.58
−0.46 =
1.04
and
x = 3.58 + (-0.46)(1.04) = 3.10
Thus 67.72% of the daily average amount of solid waste per person weighs more than 3.10
pounds.
EXERCISE. Solve the following.
1. Determine the probabilities for the following normal distribution problems.
a. = 604, = 56.8, x 635
b. = 48, = 12, x < 20
c. = 111, = 33.8, 100 x < 150
d. = 264, = 10.9, 250 < x < 255
e. = 37, = 4.35, x > 35
2. On a statistics examination the mean was 78 and the standard deviation was 10.
(a) Determine the standard scores of two students whose grades were 93 and 62, respectively.
(b) Determine the grades of two students whose standard scores were -0.6 and 1.2, respectively.
3. Find (a) the mean, (b) the standard deviation on an examination in which grades of 70 and 88
correspond to standard scores of -0.6 and 1.4, respectively.
4. Find the area under the normal curve between
(a) z = -1.20 and z = 2.40
(b) z = 1.23 and z = 1.87
(c) z = -2.35 and z = -0.50.
Vicente G. Padilla
Associate Professor IV, CatSU College of Science
5. Find the area under the normal curve
(a) to the left of z = -1.78
(b) to the left of z = 0.56
(c) to the right of z = -1.45
(d) corresponding to z 2.16
(e) corresponding to -0.80 z 1.53
(f) to the left of z = -2.52 and to the right of z = 1.83.
6. Find the values of z such that
(a) the area to the right of z is 0.2266
(b) the area to the left of z is 0.0314
(c) the area between -0.23 and z is 0.5722
(d) the area between 1.15 and z is 0.0730
(e) the area between –z and z is 0.9000.
7. The mean grade on a final examination was 72, and the standard deviation was 9. The top 10%
of the students are to receive A’s. What is the minimum grade a student must get in order to
receive an A?
8. The heights of 300 students are normally distributed with mean 69.0 inches and standard
deviation 3.0 inches, how many students have heights
(a) greater than 72 inches
(b) less than or equal to 64 inches
(c) between 65 and 71 inches inclusive
(d) equal to 68 inches?
Assume the measurements to be recorded to the nearest inch.
Vicente G. Padilla
Associate Professor IV, CatSU College of Science