EXAMPLE
The 50 companies’ percentages of revenues spent on R&D (i.e. Research and Development) are:
13. 6. 13.
9.5 8.2 6.5 8.4 8.1 7.5 10.5
5 9 5
7.2 7.1 9.0 9.9 8.2 13.2 9.2 6.9 9.6 7.7
9.7 7.5 7.2 5.9 6.6 11.1 8.8 5.2 10.6 8.2
11. 7.
5.6 10.1 8.0 8.5 11.7 7.7 9.4 6.0
3 1
8.0 7.4 10.5 7.8 7.9 6.5 6.9 6.5 6.8 9.5
Calculate the proportions of these
measurements that lie within the intervals , , and , and Compare the results with
the theoretical values. The mean and standard deviation of these data come out to be 8.49 and 1.98,
respectively.
Hence
= (8.49 – 1.98, 8.49 + 1.98)
= (6.51, 10.47)
Now count the values between this range.
A check of the measurement reveals that 34 of the 50 measurements, or 68%, fall between 6.51and 10.47.
Similarly, the interval
= (8.49 – 3.96, 8.49 + 3.96)
= (4.53, 12.45)
Now count the values between this range.
Contains 47 of the 50 measurements, i.e. 94% of the data-values
Finally, the 3-standard deviation interval around , i.e.
= (8.49 – 5.94, 8.49 + 5.94)
= (2.55, 14.43)
Now count the values between this range.
Contains all, or 100%, of measurements.
In spite of the fact that the distribution of these data is skewed to the right, the percentages of data-values
falling within
1, 2, and 3 standard deviations of the mean are remarkably close to the
Theoretical values (68%, 95%, and 100%)
Given by the Empirical Rule.
In this example, 94% of the values are lying inside the interval , and this percentage is greater
than 75%.
Similarly, 100% of the values are lying inside the interval , and this percentage is greater than
89%.
Expressing these relationships in terms of the standard deviation, which measures dispersion, we can say
that when the values of a set of data are concentrated near their mean, the standard deviation is small. And
when the values of a set of data are scattered widely about the mean, the standard deviation is large. In
exactly the same way, if the standard deviation computed from a set of data is large, the values from
which it is computed are dispersed widely about their mean.
EXAMPLE
Suppose that a study is being conducted regarding the annual costs incurred by students
attending public versus private colleges and universities in the United States of America. In
particular, suppose, for exploratory purposes, our sample consists of 10 Universities whose
athletic programs are members of the ‘Big Ten’ Conference? The annual costs incurred for
tuition fees, room, and board at 10 schools belonging to Big Ten Conference are given as
follows:
Name of University Annual Cost
(in$000)
Indiana University 15.6
Michigan State University 17.0
Ohio State University 15.2
Pennsylvania State University 16.4
Purdue University 15.2
University of Illinois 15.4
University of Iowa 13.0
University of Michigan 23.1
University of Minnesota 12.3
University of Wisconsin 14.9
If we wish to state the five-number summary for these data, the first step will be to arrange our
data-set in ascending order:
Ordered Array:
X0=13.0 14.3 14.9 15.2 15.2 15.4 15.6 16.4 17.0 Xm=23.1
And if we carry out the relevant computations, we find that:
• The median for this data comes out to be 15.30 thousand dollars.
• The first quartile comes out to be 14.90 thousand dollars, and
X0=13.0 14.3 14.9 15.2 15.2 15.4 15.6 16.4 17.0 Xm=23.1
• The third quartile comes out to be 16.40 thousand dollars.
Therefore, the five-number summary for this data-set is:
The Five-Number Summary:
13.0 14.9 15.3 16.4 23.1
If we apply the rules that I conveyed to you a short while ago, the annual cost data for our
sample are right-skewed. We come to this conclusion because of two reasons:
• The distance from Q3 to Xm (i.e., 6.7) greatly exceeds the distance from X0 to Q1 (i.e., 1.9).
• If we compare the median (which is 15.3), the mid-quartile range (which is 15.65), and the
midrange (which is 18.05), we observe that the median < the mid-quartile range < the midrange.
Both these points clearly indicate that our distribution is positively skewed.