Module 4 Descriptive Statistics for a Single Variable
Module 4 Descriptive Statistics for a Single Variable
Data Examples
Those in retail, for example, might analyze information gathered on purchases at their company
throughout the year, to make sure they have the right amount of product at the right time. By
keeping what is needed on the shelves, a retailer will boost sales. By eliminating product that is not
selling, retailers free up shelf space for other, more successful products.
Those in manufacturing might look at worker productivity on different machines in order to make
good purchasing decisions. Those in a service industry might use a survey to capture which
aspects of their service customers were happy with, and which need improvement.
The information described above is available as raw data (number of products purchased at a
particular time or ratings of a service). But to make use of the information, you need to understand
how to work with it. Perhaps most important, you need to be able to decide when and if the data is
really giving you actionable information.
For example, suppose in a month, one of the tax consultants you employ achieves an average
rating of 4. 2 (out of five) from the six customers who reviewed him out of the 70 he served. A
second employee received an average rating of 3. 5 (out of five) from twenty of the100 she served.
The second employee seems more productive, but less popular. But how significant are these
ratings? Can we say definitively that the second employee is less effective at meeting her
customers' expectations? Did the lower work load of the first employee allow him to better meet
customer expectations?
Only when you've correctly analyzed the data will you be able to answer the questions above. If
you can answer them, you can be comfortable with decisions about whom to employ and the
optimal workload to assign.
Business leaders will need to gather and analyze data to answer a wide range of questions in order
to make sound business decisions.
Types of Data
Before we discuss different types of graphs and how to interpret them, we need to understand the
types of data that we will be representing in graphs.
1. Quantitative data, also called numerical data, consists of data values that are numerical,
representing quantities that can be counted or measured.
2. Categorical data, also called qualitative data, consists of data that are groups, such as
names or labels, and are not necessarily numerical.
Note: It is possible for numbers to be used as categorical data. For example, the numbers on
the uniforms of basketball players are categorical, because they are used to identify players, and
they do not measure a quantity. Zip codes are another example of numbers that do not measure
quantity but are used to categorize different locations by the postal system.
Examples
Quantitative (numerical) examples would
include the number of employees at your
firm or the average salary of an IT
professional.
It is part of human nature to learn about and consume the knowledge around us, and graphical
displays are a tool that helps us to do that. Graphs of data sets can help us better understand,
organize, and present data and information in a simple, visual way. This module will include
measurements and graphical displays only for single variable data. The next module will discuss
graphical displays for two variables.
The type of graph used to display data is dependent upon whether the data is categorical or
quantitative. First, we will look at two different kinds of graphical displays that are used to illustrate
categorical data. The style of graph, as well as how to describe a distribution from their display, will
be discussed in further detail on the following pages.
The point of a visual display of data is to help your audience grasp the
implications of the numbers, so it's important to know when and how to use
different visual displays.
Two of the most common graphical displays for categorical data are pie
charts and bar charts. Pie charts provide a vivid sense of how different
components make up a whole, while bar graphs give a good sense of
change across categories or time.
For example, suppose the customer base for your ice cream company is 64% women, and you
think there's an untapped male market out there. You might want to emphasize the extent to which
your product or service is failing to appeal to men by creating a pie chart.
A bar graph is a wonderful way to compare different categories or data across time. You might
chart yearly sales of your various ice cream flavors to show your colleagues how poorly a particular
flavor faired. A bar graph might help your argument that this flavor needs to be retired by showing
just how little of it sold.
All graphs can be used to mislead, but pie charts have a particular weakness that you should be
aware of. They can be difficult for viewers to interpret if there are a number of categories that all
have similar proportions. Most viewers will not be able to easily distinguish 20 % from 27 % if there
are 4 or 5 categories ranging from 10 % to 30 % each. If you want to emphasize the differences
among the categories, a bar graph is a better choice.
To create a pie chart, the percentage of the whole that each category represents must be calculated
from the raw data. Consider the following data distribution from a sample population representing
the number of hours exercised per week.
Using the percentages calculated from the raw data we can now create a pie chart.
Rather than pieces of a pie, bar charts graphically illustrate data using bars. There is a bar for each
category. The height of the bar is determined by the number of values in that category. The number
of values could also be the relative frequency or the percentage. Here is an example of a bar chart
to represent the number of sales for Company XYZ by each month. The categorical variable of the
months of the year is along the horizontal axis. The vertical axis represents how many sales were
completed in that month.
Pie Charts
Below are two pie charts illustrating the customer demographic data for two different products. The
first pie chart displays the education attainment of customers who purchased Product A. The
second pie chart illustrates the education attainment of customers who purchased Product B.
The number of computer help desk visits per day of the week is shown below:
a. Monday
b. Tuesday
c. Wednesday
d. Friday
9. Refer to the bar chart above. Approximately how many computer help
desk visits occur on Wednesday?
10. Refer to the bar chart above. What is the approximate difference
between the number of computer help desk visits on Wednesday versus
Sunday? Enter your answer as a multiple of 5.
There are a variety of different ways to describe the distribution of a categorical variable. Consider
the following bar chart that illustrates a distribution of blood types and Rh factor:
The bar chart shows the frequency distribution for the categorical variable.
In the above chart, the categorical variable is Blood Type
We can make some additional observations about the data based on these displays:
1. The most common type of blood is Type O+ (38 % ), followed by Type A+ (34 % ), Type B+ (
9 % ), and Type O- (7 % )
Career Connections
For example, the cost to maintain some operating systems is more than
others. Another important consideration is the level of tech support
available. What does all of this have to do with statistics, though? If you're
managing many computers, some of which may be on different operating
systems, you may need a quick summary of the different types of operating
systems and how many there are of each kind in the company or the team you work in. A quick
way to get information like this is a bar chart. You can then use the information in the bar chart to
forecast what resources (time and/or money) each operating system will require to stay
operational.
In short, a bar charts is a great way to summarize categorical data (data that falls into distinct
categories) such as the operating systems a group of computers use. Other examples you might
see in the IT field:
The available retirement options and how many people participate in each one
The number of full-time, part-time, and temp positions in your organization
In short, bar charts occur in many different disciplines and they all give you a quick way to
summarize a single categorical variable in an effective and efficient manner.
Exercise
3. True or False: Based on the data above, an occupational injury fatality is more likely to occur
from exposure to harmful substances and environments than from assaults and violent acts.
4. True or False: For this data set the categorical variable is the reason for occupational injury
fatality.
The table below lists various kinds of graphical displays used for quantitative data. The type of
graph, as well as how to describe the data distribution from each, will be discussed in further detail
on the following pages.
Career Connections
Dot plots and stem plots are useful for getting a visual sense of a small set
of data. For example, suppose you are considering offering a pension for
the 20 employees of your small business. You might chart the employees'
ages to get a better grasp on whether there are any large clumps of
employees who might retire at similar times. Or if you have an employee
you suspect is under-performing on the job, you might randomly select 15
other employees and chart their productivity (in terms of customers served)
with that of the suspected employee to see if his falls far below or in line with the others'. Both of
these exercises would give you a good sense of what the data is telling you.
A dot plot shows each data value as a point, distributed along a horizontal axis. Dot plots are useful
because they show the distribution of a data set, as every data value is represented by a dot.
Example
The table below shows the number of computer lab visits each day for20 days. Construct a dot
plot for the data.
43 51 13 31 20
32 32 23 32 2
52 57 44 36 33
44 32 45 14 25
Step 1
2, 13, 14, 20, 23, 25, 31, 32, 32, 32, 32, 33, 36, 43, 44, 44, 45, 51, 52, 57
Step 2
Step 3
Place a dot above the horizontal axis for each data point in the table. Here, each "dot" is
represented by the letter x. Any repeated values, (such as 32 or 44 which have red boxes around
them), should be represented by a mark for each value, stacked vertically.
Stem plots
Stem plots, also called stem-and-leaf plots, are another way to show a data set and its distribution
or shape. A stem plot is constructed by separating each data value into a stem (usually the left-most
digit) and a leaf (usually the right-most digit). For example, if 36 is a data point, the 3 would be the
stem and the 6 would be the leaf. The data is arranged in two columns, stems and leaves, with a
vertical line separating the columns.
Example
Using the same data from above, a stem-and-leaf plot would look like this:
Stems Leaves
0 2
1 34
2 035
3 1222236
4 3445
5 127
Exercise
1. Refer to the dot plot above. What is the minimum value in the data
set?
2. Refer to the dot plot above. What is the maximum value in this
data set?
3. Refer to the dot plot above. Is the value of 0 in the data set?
(Enter Yes or No)
4. Refer to the dot plot above. During this network test how many
reports recorded a download speed of 60 Mbps?
5. Refer to the dot plot above. What download speed was most often
recorded during this network test?
Stem Plots
The following stem plot will be used for Questions 6-10:
4.04.2 Histograms
Histograms
Career Connections
In this lesson, you'll learn about working with histograms from many different disciplines, such as
business and healthcare. In the end, histograms can be used to help us understand a single
quantitative variable better. Be sure to think about how you could use histograms in your own
discipline as you go through this material.
A histogram is a graph that displays quantitative data. The vertical bars in a histogram show the
counts or numbers in each interval. A comparison of the intervals, or a review of the graph as a
whole, helps the audience understand the information presented.
The distinction between a histogram and a bar chart is an important distinction to make. As
previously discussed a bar chart measures categorical data that is distributed over groups or
categories, while a histogram measures how quantitative data is distributed over various intervals.
For example, a histogram would be appropriate to display how many people fall in various intervals
of heights, as height is an example of quantitative data.
In other words, a histogram is used to display frequencies or relative frequencies for quantitative
data; in contrast, a bar chart is used to display frequencies (i.e., counts) or relative frequencies for
categorical data.
Bar Chart: displays frequencies (i.e., counts) or relative frequencies for categorical data
Histograms allow team members and stakeholders to view a significant amount of data at one time,
and to see how data is distributed across various intervals of values. The histogram's bars
represent the values or intervals in the study. The height of each bar shows how many
observations or events fall into each interval. The shape of the graph illustrates how the data is
distributed.
The intervals in a histogram need to encompass all of the data collected. It is important to make
sure the minimum and maximum values are accounted for, as well as every value in between.
Additionally, make sure that the intervals are comparable and exhaustive. The intervals should run
consecutively so that all data is accounted for and the visual representation of the graph is
accurate.
It is also important to use an appropriate number of intervals; if you have difficulty determining how
many intervals to use, you can refer to the rough guidelines laid out in the chart below:
This symmetry, or type of distribution, is illustrated in the histogram below. Notice how the
middle value(s) is the most frequent. The values decrease in a symmetrical manner, to the
right and left of the center of the histogram.
A bell curve, or normal distribution has a very specific distribution. Normal distributions
will be discussed later in this module. For now, it is important to note that just because a
histogram is symmetric does not make it normal!
Skewed Distribution
Skewed distribution
Distributions can also be asymmetric. A skewed distribution is the term used to describe a
distribution that has a "long tail" on one side of the peak. In other words, the distribution is
lopsided and does not have a symmetrical shape. Skewness is used to measure the
asymmetry of a distribution. When more data falls further to the left of the peak, it is known
as skewed left. This type of distribution is also referred to as negatively skewed.
A histogram that is skewed right, or positively skewed indicates that more values are far
greater than the most common value, but not far less. A histogram that is skewed left, or
negatively skewed, indicates that more values are far less than the most common value,
but not far greater.
Career Connections
Histograms
As with other graphs, histograms and bar charts help people grasp and interpret data. Therefore,
the charts can be used to organize data to see if recognizable patterns emerge. For example, if you
When you plot the number of items sold at a certain price points on a histogram, you might notice
that they cluster at $ 20 to $ 29 and $ 80 to $ 89. More expensive items, originally priced at $ 40 to
$ 50 did not sell until they were discounted into the$ 20 to $ 30 range while merchandise originally
priced at $ 80 to $ 90 (as well as more expensive merchandise discounted to that range) sold
relatively well. Using this information, you might design or price more items to fall into the popular
ranges, hoping to capitalize on your customers' comfort level with these price points.
While a tabular record of this data would convey the same information if carefully inspected, the
height of the bars at $ 20 to $ 29 and $ 80 to $ 89 leaps out at the viewer. The relative differences of
each bar can be more easily grasped than a set of numeric data.
Different distributions hold different shapes. As we have seen earlier a distribution can be
symmetric:
There are also other common shapes distributions can hold. As well as bell-shaped distributions,
another common symmetric distribution is U-shaped. A U-shaped distribution occurs when a
symmetric distribution has a "valley" rather than a peak:
If a distribution has two clear peaks rather than one, it is known as bimodal. A histogram with two or
more clear peaks is called multimodal. Below is an example of a histogram that is bimodal.
When describing quantitative data, center and spread are two important characteristics. The center
of a set of quantitative data is a point that represents the "middle" of the data. As we will see, this
can be measured in many different ways. There are also multiple measurements that are used to
describe spread. Spread, in general, is a way to describe the dispersion of quantitative data. Is all of
the data clustered around one point, or is it spread out?
We can create a graphical display in order to better understand the shape, center, and spread of the
data. A histogram would be a good choice.
Career Connections
Exercise
Exercises: Distribution
Answer the question below by analyzing the following histogram. Enter the letter that corresponds
with your answer.
A. Skewed left
B. Skewed right
C. Uniformly distributed
D. Normally distributed
2. Data that is skewed right is:
A. Negatively skewed
B. Positively skewed
C. Symmetric
D. Normally distributed
3. Which of the following does not describe a symmetric histogram?
A. Bell-shaped
B. U-shaped
C. Positively Skewed
D. Uniform
Answer the questions 4 and 5 below by analyzing the following histogram. Enter the letter that
corresponds with your answer.
A. Bimodal
B. Skewed left
C. Uniform
D. Skewed right
5. What can we infer about this data?
A. Bimodal
B. Skewed left
C. Skewed right
D. Symmetric
7. A histogram with a symmetric distribution that has a "valley" rather
than a peak is described as:
A. Unimodal
B. U-shaped
C. Uniform
D. Bell shaped
A. Negatively skewed
B. Positively skewed
C. Symmetric
D. Normally distributed
10. A histogram that has more than two modes is known as:
A. Unimodal
B. Uniformly distributed
C. Multimodal
D. Symmetric
A data set is any collection of numerical values, such as measurements, observations, or survey
responses. For example, if we measure the heights (in centimeters) of ten randomly selected
people, we could have the following data set:
In statistics, measurements need to be both reliable and valid. Reliable data is both consistent and
repeatable. If you were to administer the same test to the same person three times and the scores
were similar each time, the test could be categorized as reliable. If the results varied greatly, the
test would be unreliable. Similarly, valid data is data resulting from a test that accurately measures
what it is intended to measure. For instance, if a test reflects an accurate measurement of a
student's abilities, it is said to be valid.
Communicating Calculations
Previously, we've roughly estimated the center, spread, and possible outliers of data. For more
precise results, we calculate these measures.
Reliable and valid data calculations give rise to more accurate and precise results. This precision
allows for additional graphical displays for quantitative data. Rather than rough estimations, we can
display and report precise figures that measure spread, center, and other summary values of data.
Mean
The mean is one of the most useful measures of central tendency. The mean, also known
as the average, is a single value that represents the center of a set of data values. Mean
can be substantially influenced by one or more extreme values in a data set (think skewed
data), so mean is only used when the data is symmetric. Therefore, we say that the mean
is not a resistant measure of center.
To calculate the mean, the values in a data set are simply added together and divided by
the number of available values.
Let's use the ten heights from the example on the previous page.
Step 1
To find the mean, the first step is to add all of the values together.
Step 2
Next, you divide that sum (1769.1) by the number of values in the data set 10).
(
1769. 1 ÷ 10 = 176. 91
Example
Mr. Nolan coaches a little league baseball team made up of 18 players. On the team, there
are 48-year-olds, 27-year-olds, 36-year-olds, 39-year-olds, 210-year-olds, 35-year-olds, and
112-year-old. Eleven of the players are boys and seven are girls. What is the mean age on
Mr. Nolan's baseball team?
To calculate the mean, sort out the data that's applicable to the team ages. Though within
the problem those of the same age are grouped together, to find the mean, we need to
account for each of the 18 players' ages individually. The gender of the players is not
relevant in this problem.
8 + 8 + 8 + 8 + 7 + 7 + 6 + 6 + 6 + 9 + 9 + 9 + 10 + 10 +
5 + 5 + 5 + 12 = 138
138 ÷ 18 = 7. 6666. . .
Median
The second measure of central tendency is the median. The median is the "halfway" point
of a set of values; an equal number of values will fall above and below the median of a
data set.
Unlike the mean, the median is not overly influenced by extreme values in the data set, so
we can use the median when the data is skewed. Therefore, we say that the median is a
resistant measure of center. To properly find the median, values must be first sorted from
smallest to largest.
Step 1
Step 2
Next, count the total number of values in the data set to find the median. If there is anodd
number of values, the halfway point will fall directly on a value and that will be your
median.
There are 15
values in total, which means the median, or halfway point of the data set, is 72
. (There are seven values below 72
, and seven values above 72
.)
This median can be calculated by adding those two middle values and dividing by two.
= (145) ÷ 2
= 72. 5
Step 1
Step 2
Imagine drawing a line down the middle of your data so that half the data points are on the
right of the line and half are on the left.
If your line lands on a data point (for example, if you have an odd number of values), that
is the median.
If your line lands in between two data points (for example, if you have an even number of
values), your median is the halfway point between the two data points, and the median is
calculated by taking the average of those numbers. In this case:
Mode
The mode is the third and final measurement of central tendency. The mode represents the
value that occurs most often in a data set.
The mode is only relevant if a data set has values that are repeated, and unlike mean and
median, there can be more than one mode in a data set.
Example
Mode is a measurement easily displayed in graphs -
also unlike mean and median. Take a look at the histogram below, which presents a range
of test scores, and decipher which bar in the graph displays the mode.
1. In fact, if all of the intervals contained the same number of scores, every value would
be the mode, rendering that measure useless. For this reason, you should always
graph your data. The other two measures of central tendency, median and mean,
may not tell you this information about a data set's distribution as well as a graph.
Career Connections
Career Connections
What do these two data sets tell you? First, it is reasonable to expect that any employee should be
able to clip toenails in 11
minutes. When you create the employee schedules, you know pretty precisely how much time to
allow. Dog washing, however, requires further investigation. What makes the difference in the
times? In a new data set, you might chart the time taken to wash the dog against its weight (a
proxy for size) or against its manageability. When you have identified the cause of the spread, you
can ask for the relevant information from the customer and use it to create efficient schedules for
your dog groomers. Until you have a better grasp of how to predict the amount of time needed for
dog washing, you will have difficulty creating a schedule.
Range
Range
Range is the difference between the smallest (minimum) and greatest (maximum) values of
a data set.
The minimum is the smallest value available in a data set. The maximum is just the
opposite; it's the greatest value in a data set.
Example
18, 11, 3, 26, 13, 40, 31, 5, 12, 45, 52, 22, 17, 33, 8
To find the minimum, maximum, and range, it is helpful to first sort the values from smallest
to largest.
3, 5, 8, 11, 12, 13, 17, 18, 22, 26, 31, 33, 40, 45, 52
After the data is sorted, the minimum and maximum fall at either end of the data set. In the
data set shown, the minimum is 3
and the maximum is 52
.
To find the range, simply subtract the minimum from the maximum.
Range =
maximum -
minimum
52 - 3 = 49
Interquartile Range
Quartiles are widely used measures when dealing with data sets. Quartiles are values that
divide a data set into four equally sized groups. There is one median per dataset that splits
the data into two equally sized groups. Similarly, a dataset has three quartiles that split the
data into four equally sized groups. The interquartile range measures the difference
between the third quartile and the first quartile. To illustrate how to find the first and third
quartiles we will use the data from the following research poll that asked individuals how
many alcoholic drinks they have per week.
Example
Polling 12
people about how many alcoholic drinks they have per week might yield the following
data:
9, 1, 7, 5, 4, 3, 6, 2, 9, 1, 9, 4
1, 1, 2, 3, 4, 4, 5, 6, 7, 9, 9, 9
As we have 12
values in our data set, placing the data into the four quarters results in having three data
values in each quarter. As you can see in the chart above, specific numbers that have
multiple occurrences are included the number of times they occur.
The interquartile range is an indicator of the distribution of a sample and can also help
identify any outliers. Outliers are data points (numbers) that are far away from all other data
points. It is helpful to identify any outliers and determine whether they should be used.
1, 1, 2, 3, 4, 4, 5, 6, 7, 9, 9, 9
2. Find the median, or midpoint, of the data set. This can also be called the second
quartile (Q2
).
1, 1, 2, 3, 4, 4
| 5, 6, 7, 9, 9, 9
Q2 = 4. 5
3. Identify the median of the lower half of the data set and label it asQ1
(the first quartile).
1, 1, 2
In this case, the median of the lower half of the data set is midway between2
and 3
, which averages to 2. 5
.
Q1 = 2. 5
4. Identify the median of the upper half of the data set and label it asQ3
(the third quartile).
1, 1, 2
| 3, 4, 4
| 5, 6, 7
| 9, 9, 9
In this case, the median of the upper half of the data set is midway between7
and 9
, which averages to 8
.
Q3 = 8
5. Subtract Q1
from Q3
to determine the interquartile range, or IQR.
1, 1, 2,
| 3, 4, 4,
| 5, 6, 7,
| 9, 9, 9
IQR = Q3
-
Q1
Q1 = 2. 5
IQR
= 8 - 2. 5 = 5. 5
IQR
= 5. 5
Let's look at one more example of finding the IQR for a dataset.
Example
Find the IQR of the following set of download speeds from various clients of a certain
ISP.
21, 8, 41, 20, 11, 16, 13, 14, 35, 27, 40, 18, 30, 32, 20
Follow these steps to find the interquartile range of a data set with an odd number of data
points.
8, 11, 13, 14, 16, 18, 20, 20, 21, 27, 30, 32, 35, 40, 41
Q2 = 20
3. Identify the median of the lower half of the data set and label it asQ1
(the first quartile). In this case, there are 7
data points in the lower half of the data, so Q1
will be the 4th
in that list:
8, 11, 13, 14, 16, 18, 20, 20, 21, 27, 30, 32, 35, 40, 41
Q1 = 14
4. Identify the median of the upper half of the data set and label it asQ3
(the third quartile). In this case, there are 7
data points in the upper half of the data, so Q3
8, 11, 13, 14, 16, 18, 20, 20, 21, 27, 30, 32, 35, 40, 41
Q3 = 32
5. Subtract Q1
from Q3
to determine the interquartile range, or IQR.
IQR = Q3
-
Q1
Q3 = 32
Q1 = 14
IQR
= 32 - 14 = 18
IQR
= 18
1. Recall the data set from the previous example, placed in order from least to
greatest.
1, 1, 2, 3, 4, 4, 5, 6, 7, 9, 9, 9
Q3 = 8
Q1 = 2. 5
IQR =
Q3
-
Q1 = 8 - 2. 5 = 5. 5
IQR = 5. 5
IQR × 1. 5 =
Q3 = 8
8 + 8. 25 = 16. 25
Q1 = 2. 5
2. 5 - 8. 25 = - 5. 75
1, 1, 2, 3, 4, 4, 5, 6, 7, 9, 9, 9
Five-number Summary
The five-number summary lists the minimum, first quartile, median, third quartile, and
maximum in a data set. The five number summary will be represented by a graphical
display that we will learn about later in Module 4 called a box plot.
The following table is a five-number summary of the hours spent volunteering in the past
year taken from a sample population of employees.
Five-number Summary
Statistic Data Value
The standard deviation tells you how far, on average, the data points are from the mean. In
this course, we will not focus on how to calculate the standard deviation (we can use
computers to do this for us) but rather on building an intuition for using the standard
deviation to measure how spread out the data is in a dataset. Standard deviation is a
measurement that is used for symmetric data.
Approximately 68 %
of all values are within 1
standard deviation of the mean
Approximately 95 %
of all values are within 2
From the Standard Deviation Rule, we can calculate all of the parts of the bell-curve. It is
important to memorize the Standard Deviation Rule. You should also know how to
calculate the other percentages, or you can memorize those.
Let's take a look at a couple examples to see how we can apply the Standard Deviation
Rule.
Example
Custom log home builder, Tinker Log Homes Inc. takes an average length of40
weeks, or 280
days, with a standard deviation of 13
days to build a custom log home (assume a normal distribution). Knowing this information,
and using the Standard Deviation Rule, let's answer the following questions.
1.
68 % of the data is between what two values?
The curve above illustrates this distribution. Using the Standard Deviation Rule, we
know that 34 %
of the data is one standard deviation above the mean (+1
SD), and 34 %
of the data is one standard deviation below the mean (-1
SD). Therefore, according to the Standard Deviation Rule, 68 %
Answer: 68 %
of the values will fall between 267
and 293
days.
2. What percentage of custom log home build timelines will range between 254
and 306
days?
To answer this question, first, we have to figure out how far both254
and 306
are from the mean.
The difference between both of these values and the mean is equal to26
. Next, we divide this value by our standard deviation of13
. The result is equal to 2
. Therefore, 254
and 306
Now that we know that these two values are two standard deviations from the mean
of 280
days, we can use the Standard Deviation Rule to determine what percentage of
home build timelines will last between 254
and 306
days.
As illustrated in the distribution curve above, the Standard Deviation Rule states that
13. 5 %
of the data falls between one and two standard deviations from the mean. Knowing
this information, we can now calculate the percentage of data that falls between two
standard deviations (or between 254
and 306
days) from the mean by adding up the percentage values in the shaded areas of the
curve.
3. What is the chance that a custom log home build will last longer than306
days?
The percentage values under the curve that are greater than two standard
deviations from the mean are illustrated in the distribution curve below.
We refer to values in a histogram that come after a gap as extreme values and possible outliers.
Due to the fact that this distribution skews to the right, the most extreme values are on the right side
of the distribution. These extreme values, or possible outliers, have an effect on the measures of
center.
Imagine you are looking at income data among a group of people. Most of the people in this data
set are making between $30, 000
and $75, 000
per year. There are a few high-earners who have incomes over $100, 000
and there is one person who is earning $10, 000, 000/
year. This distribution is skewed-right. The few individuals with a very high income skew the mean
to the right. You may find that the mode income is $40, 000
, the median income is $55, 000
, but the mean income is $100, 000
. In this sample, $40, 000
is the most common income, and the middle-earner makes $55, 000/
year. The mean income is less useful, though, as the highest earners have dramatically skewed the
mean to the right.
The following table summarizes the preferred measures of center and the measures of spread for
normal and skewed distributions.
Measures of Measures of
Distribution
Center Spread
Skewed Median Range or IQR
Normal Symmetric Mean Standard Deviation
A modified box plot is just like a regular box plot, except that outliers are shown as points above or
below the minimum and maximum. The box plot below illustrates an outlier. In this example, the
outlier is beyond the end of the whisker.
Five-number Summary
A box plot is a convenient way to show five important statistical values: minimum, maximum, first
quartile, median, and third quartile. As previously mentioned, these five values are often referred to
as a five-number summary of a data set. Many statistical outputs from technology, such as
calculators, computer programs, or other software, have five-number summary displays built in.
Five-number Summary
The far left side of the box plots represents the minimum value of the data set. In the
above example, this is 60
.
The left whisker represents the lower 25 %
of the data. In this example, that is from 60
to 70
. That means 25 %
of students scored between 60
and 70
on the algebra exam.
The first part of the box represents the next 25 %
of the data. It is shaded blue in the above box plot. So another 25 %
of the students scored between the Q1
of 70
and the median of 75
.
The line in the middle of the box is the median Q(2
). That means 50 %
of the data is below this point, and 50 %
of the data is above this point. In the above example, 50 %
Five-number Summary
The top of the line is the maximum value in the data set; in the above example, the
Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.
The top of the line is the maximum value in the data set; in the above example, the
maximum value is approximately 18
.
The upper limit of the rectangular box, shaded in purple in the above example, represents
the third quartile (Q3
).
The line in the middle of the rectangle (in the above example, this line is indicated where
the green and purple shaded areas of the rectangle meet) is the second quartile (Q2
) or median.
The lower limit of the rectangular box (in the above example, shaded in green) represents
the first quartile (Q1
).
The bottom of the line represents the minimum value of the data set; in the above
example, this is approximately 7. 5
.
Exercise
Box Plots
The box plot below shows the distribution of the time a real estate agent spends with a client prior
to the client closing on a new home. Answer the questions below by analyzing the box plot.
Minimum:
Q1
:
Median:
Q3
:
Maximum:
7.
Identify the Five-number Summary values by entering in the letter on the graph that corresponds to
that value.
Minimum Value:
Q3
:
Maximum Value:
Q1
Review Checkpoint
To test your understanding of the content presented in this assignment, please click on the
Question icon below. Click your selected response to see feedback displayed below it. If you have
trouble answering, you are always free to return to this or any assignment to re-read the material.
1. Which of the following represent the five values included in the five-number summary of a data
set?
Correct.
2. Based on the following box plot, what percent of initial client consultations last less than70
minutes?
a. 75 %
Correct. The answer is a. We know that each part of the box plot represents25 %
of the data and that 70
b. 50 %
c. 40 %
c. 25 %
3. Based on the box plot below, approximately what percent of salespeople represented in this data
set have a total first quarter sales level less than 120
?
a. 10 %
b. 15 %
c. 25 %
d. Cannot be determined
4. Based on the box plot below, what percent of the class scored between75
and 85
on the algebra exam?
a. 25 %
Correct. From this box plot we see the median value is equal to75
, and Q3
is equal to 85
. Since we know in a box plot that25 %
of the data falls between the median and Q3
, we can determine that 25 %
of the class scored between 75
b. 10 %
c. 50 %
d. 15 %
5. Based on the box plot below, what percent of the data falls between83
and 88
?
a. 5 %
b. 25 %
c. 35 %
b. 50 %
Below are just a few of the ways that graphical displays can misrepresented.
This graph shows the number of admissions per year for three universities. Do you notice
anything wrong with this graph? Compare it to the following:
In the first graph, the differences between the universities' admissions appear to be greater
than they do in the second graph. The reason is that the vertical scale does not start at
zero. This is known as truncating, which exaggerates the differences between the three
universities.
Another example of graphs that can misrepresent data are graphs that omit labels or units,
such as the pie chart displayed below.
4.11 Flashcards
Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.
4.11 Flashcards
Module 4 Flashcards
Term Definition