0% found this document useful (0 votes)
3 views72 pages

Module 4 Descriptive Statistics for a Single Variable

Module 4 focuses on descriptive statistics for a single variable, teaching how to apply statistical measures and graphical displays to analyze data. It covers types of data, including quantitative and categorical, and emphasizes the importance of correctly interpreting data for business decisions. The module also discusses various graphical displays, such as pie charts and bar charts, and their appropriate use in representing different types of data.

Uploaded by

phillylpm
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views72 pages

Module 4 Descriptive Statistics for a Single Variable

Module 4 focuses on descriptive statistics for a single variable, teaching how to apply statistical measures and graphical displays to analyze data. It covers types of data, including quantitative and categorical, and emphasizes the importance of correctly interpreting data for business decisions. The module also discusses various graphical displays, such as pie charts and bar charts, and their appropriate use in representing different types of data.

Uploaded by

phillylpm
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Module 4: Descriptive Statistics for a Single Variable

Module 4: Descriptive Statistics for a Single Variable

4.01 Learning Objectives


Module 4: Learning Objectives
After completing this module, you should be able to:

1. Apply the standard deviation rule to a special case of normal distributions


2. Identify the appropriate numerical measures for a given set of data in different contexts
3. Relate measures of center and spread to the shape of the distribution
4. Calculate the mean, median, mode, quartiles, range, and interquartile range for a set of
quantitative data
5. Evaluate data from several different graphical displays of a distribution of a categorical
variable (bar chart, pie chart) and of a quantitative variable (histogram, stem plot, box plot)
6. Describe the distribution of a categorical variable in context
7. Describe the distribution of a quantitative variable in context: a) describe the overall pattern,
b) describe striking deviations from the pattern
8. Identify the appropriate graphical display for a given set of data in different contexts
9. Explain how graphical displays can be used to misrepresent data

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


4.02 Types of Data
Career Connections

Data Examples

Because of the relatively low cost of computer storage and processing


power, it is now possible to collect and store vast amounts of data.
However, a large set of data in itself is not very valuable. If it's correctly
interpreted, on the other hand, it can enable smart, effective business
decisions. Business leaders need to understand statistics to make the best
decisions for their companies.

Those in retail, for example, might analyze information gathered on purchases at their company
throughout the year, to make sure they have the right amount of product at the right time. By
keeping what is needed on the shelves, a retailer will boost sales. By eliminating product that is not
selling, retailers free up shelf space for other, more successful products.

Those in manufacturing might look at worker productivity on different machines in order to make
good purchasing decisions. Those in a service industry might use a survey to capture which
aspects of their service customers were happy with, and which need improvement.

The information described above is available as raw data (number of products purchased at a
particular time or ratings of a service). But to make use of the information, you need to understand
how to work with it. Perhaps most important, you need to be able to decide when and if the data is
really giving you actionable information.

For example, suppose in a month, one of the tax consultants you employ achieves an average
rating of 4. 2 (out of five) from the six customers who reviewed him out of the 70 he served. A
second employee received an average rating of 3. 5 (out of five) from twenty of the100 she served.
The second employee seems more productive, but less popular. But how significant are these
ratings? Can we say definitively that the second employee is less effective at meeting her
customers' expectations? Did the lower work load of the first employee allow him to better meet
customer expectations?

Only when you've correctly analyzed the data will you be able to answer the questions above. If
you can answer them, you can be comfortable with decisions about whom to employ and the
optimal workload to assign.

Business leaders will need to gather and analyze data to answer a wide range of questions in order
to make sound business decisions.

Types of Data
Before we discuss different types of graphs and how to interpret them, we need to understand the
types of data that we will be representing in graphs.

Data can be divided into two types:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Two Types of Data

1. Quantitative data, also called numerical data, consists of data values that are numerical,
representing quantities that can be counted or measured.

2. Categorical data, also called qualitative data, consists of data that are groups, such as
names or labels, and are not necessarily numerical.

Note: It is possible for numbers to be used as categorical data. For example, the numbers on
the uniforms of basketball players are categorical, because they are used to identify players, and
they do not measure a quantity. Zip codes are another example of numbers that do not measure
quantity but are used to categorize different locations by the postal system.

Examples
Quantitative (numerical) examples would
include the number of employees at your
firm or the average salary of an IT
professional.

Qualitative (categorical) examples would


include the type of industry or firm (such as
finance, manufacturing, education) or
occupational titles of a group of
employees.

4.03 Graphical Displays for Categorical Data

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Graphical Displays for Categorical Data

It is part of human nature to learn about and consume the knowledge around us, and graphical
displays are a tool that helps us to do that. Graphs of data sets can help us better understand,
organize, and present data and information in a simple, visual way. This module will include
measurements and graphical displays only for single variable data. The next module will discuss
graphical displays for two variables.

Types of Graphical Displays for Categorical Data

The type of graph used to display data is dependent upon whether the data is categorical or
quantitative. First, we will look at two different kinds of graphical displays that are used to illustrate
categorical data. The style of graph, as well as how to describe a distribution from their display, will
be discussed in further detail on the following pages.

Graphical Displays for Categorical Data


Type of Display Type of Data Useful to Display
Pie Chart Categorical Different parts of a whole
Counts or frequencies for the
Bar Chart Categorical
categories

4.03.1 Pie Charts & Bar Charts


Career Connections

Pie Charts and Bar Charts

The point of a visual display of data is to help your audience grasp the
implications of the numbers, so it's important to know when and how to use
different visual displays.

Two of the most common graphical displays for categorical data are pie
charts and bar charts. Pie charts provide a vivid sense of how different
components make up a whole, while bar graphs give a good sense of
change across categories or time.

For example, suppose the customer base for your ice cream company is 64% women, and you
think there's an untapped male market out there. You might want to emphasize the extent to which
your product or service is failing to appeal to men by creating a pie chart.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


You want to move upper management to address what you perceive as an opportunity to expand
your customer base. The pie chart drives home the extent to which your company more greatly
appeals to women.

A bar graph is a wonderful way to compare different categories or data across time. You might
chart yearly sales of your various ice cream flavors to show your colleagues how poorly a particular
flavor faired. A bar graph might help your argument that this flavor needs to be retired by showing
just how little of it sold.

All graphs can be used to mislead, but pie charts have a particular weakness that you should be
aware of. They can be difficult for viewers to interpret if there are a number of categories that all
have similar proportions. Most viewers will not be able to easily distinguish 20 % from 27 % if there
are 4 or 5 categories ranging from 10 % to 30 % each. If you want to emphasize the differences
among the categories, a bar graph is a better choice.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Pie Charts
Pie charts, or circle graphs, are often used to show data as parts of a circle. A pie chart has
sections or slices. Each slice represents a category of data, and the size of each slice corresponds
to the share of the total as a percent. The sum of all the percents in a pie chart add up to 100% or
close to it because of rounding.

To create a pie chart, the percentage of the whole that each category represents must be calculated
from the raw data. Consider the following data distribution from a sample population representing
the number of hours exercised per week.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Average Number of Hours Exercised per Week
Population sample = 50
# of
Average Number of Hours People % of Sample
Raw Score ÷ Sample Size
Exercised per Week (Raw (Whole)
Score)
18
No Exercise 18 50 36 %
= 0. 36
14
1 - 2 hours 14 50 28 %
= 0. 28
6
3 - 4 hours 6 50 12 %
= 0. 12
5
4 - 5 hours 5 50 10 %
= 0. 10
4
5 - 6 hours 4 50 8%
= 0. 08
2
6 - 7 hours 2 50 4%
= 0. 04
1
More than 7 hours 1 50 2%
= 0. 02

Using the percentages calculated from the raw data we can now create a pie chart.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Bar Chart
A bar chart measures categorical data that is distributed over groups or categories. They are a
useful way to compare data among categories. For example, a bar chart would be an appropriate
graphical display to show how many people are from each state, as states are an example of
discrete categories (categorical data).

Rather than pieces of a pie, bar charts graphically illustrate data using bars. There is a bar for each
category. The height of the bar is determined by the number of values in that category. The number
of values could also be the relative frequency or the percentage. Here is an example of a bar chart
to represent the number of sales for Company XYZ by each month. The categorical variable of the
months of the year is along the horizontal axis. The vertical axis represents how many sales were
completed in that month.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Exercise

Pie Charts
Below are two pie charts illustrating the customer demographic data for two different products. The
first pie chart displays the education attainment of customers who purchased Product A. The
second pie chart illustrates the education attainment of customers who purchased Product B.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


1. What percent of customers who purchased Product A have Master's
degrees?

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


2. What percent of customers who purchased Product B have Bachelor's
degrees or higher?
3. Which is greater? (Enter the letter that corresponds with your answer)
a. The percent of customers who purchased Product A with Associate's
degrees
b. The percent of customers who purchased Product B with Associate's
degrees
4. Which group has the larger percent of "Some college, no degree?"
a. Customers who purchased Product A
b. Customers who purchased Product B
5. What is the difference in associate's degree holders between both
groups?
Bar Charts
6. A bar chart is used to show frequencies of what type of data?

a. Quantitative (numerical, continuous)

b. Categorical (qualitative, discrete)


7. For which type of data would a bar chart be most appropriate?

a. Internet download speed in Kbps

b. Total Profit by month

c. Sales volume in dollars

d. Data storage in gigabyte

The following bar chart will be used for Questions 8 to 10:

The number of computer help desk visits per day of the week is shown below:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


8. Refer to the bar chart above. Which day of the week has the greatest
number of computer help desk visits?

a. Monday

b. Tuesday

c. Wednesday

d. Friday
9. Refer to the bar chart above. Approximately how many computer help
desk visits occur on Wednesday?
10. Refer to the bar chart above. What is the approximate difference
between the number of computer help desk visits on Wednesday versus
Sunday? Enter your answer as a multiple of 5.

4.03.2 Describing the Distribution of Categorical Data


Describing the Distribution of Categorical Data

There are a variety of different ways to describe the distribution of a categorical variable. Consider
the following bar chart that illustrates a distribution of blood types and Rh factor:

Bar Chart Summary

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


By looking at the bar chart, we can summarize what we see as follows:

The bar chart shows the frequency distribution for the categorical variable.
In the above chart, the categorical variable is Blood Type

Display Data in Pie Chart

This data could be displayed in other ways, such as a pie chart:

Data: Additional Observations

We can make some additional observations about the data based on these displays:

1. The most common type of blood is Type O+ (38 % ), followed by Type A+ (34 % ), Type B+ (
9 % ), and Type O- (7 % )

2. Rh+ (84 % ) is more common than Rh- (16 % ).

Career Connections

Describing the Distribution of Operating Systems

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


There are several operating systems available in today's market. And while
Windows and iOS, the operating system for Macs, are the most popular, it's
important to keep in mind that not all operating systems are created equal.

For example, the cost to maintain some operating systems is more than
others. Another important consideration is the level of tech support
available. What does all of this have to do with statistics, though? If you're
managing many computers, some of which may be on different operating
systems, you may need a quick summary of the different types of operating
systems and how many there are of each kind in the company or the team you work in. A quick
way to get information like this is a bar chart. You can then use the information in the bar chart to
forecast what resources (time and/or money) each operating system will require to stay
operational.

In short, a bar charts is a great way to summarize categorical data (data that falls into distinct
categories) such as the operating systems a group of computers use. Other examples you might
see in the IT field:

The number of teammates who have a certain network certification


The number of teammates that have a certain kind of degree

Some examples that might be helpful across other disciplines include:

The available retirement options and how many people participate in each one
The number of full-time, part-time, and temp positions in your organization

In short, bar charts occur in many different disciplines and they all give you a quick way to
summarize a single categorical variable in an effective and efficient manner.

Exercise

Describing the Distribution of Categorical Data


1. The table below contains data regarding occupational injury fatalities for the period2004 - 2005.
Fill in the missing percentages, rounded to the nearest whole percent.

Fatal Occupational Injuries


Reason for Fatality Frequency Percent
Transportation accidents 17836
Contact with objects and equipment 7466
Assaults and violent acts 5810
Falls 4977
Exposure to harmful substances and
3317
environments
Fires and explosions 2074
Total 41480 100 %
2. Categorical variables, which represent qualitative data, are often displayed with bar charts. Use
the bar chart below and refer to your answers in the table above to answer the following questions.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Match the letter of the bar with the correct titles for each of the bars in the bar chart above.

Fatal Occupational Injuries


Bar Title
Transportation accidents
Fires and explosions
Contact with objects and equipment
Exposure to harmful substances and
environments
Assaults and violent acts
Falls

3. True or False: Based on the data above, an occupational injury fatality is more likely to occur
from exposure to harmful substances and environments than from assaults and violent acts.

4. True or False: For this data set the categorical variable is the reason for occupational injury
fatality.

4.04 Graphical Displays for Quantitative Data


Graphical Displays for Quantitative Data
Similar to displaying categorical data, there are certain types of graphs that are better suited to
illustrating quantitative data.

The table below lists various kinds of graphical displays used for quantitative data. The type of
graph, as well as how to describe the data distribution from each, will be discussed in further detail
on the following pages.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Graphical Displays for Quantitative Data
Type of Display Type of Data Useful to Display
The distribution of data, particularly
Dot Plot Quantitative clusters, gaps, and outliers. Most
useful for smaller data sets.
Stem Plot (Stem-and-leaf The distribution or shape of data
Quantitative
Plot) according to place values.
The center, spread, and outliers in a
Box Plot Quantitative
given data set.
The distribution (shape and spread) of
Histogram Quantitative
quantitative data.

4.04.1 Dot Plots and Stem Plots


Dot Plots

Career Connections

Dot Plots and Stem Plots

Dot plots and stem plots are useful for getting a visual sense of a small set
of data. For example, suppose you are considering offering a pension for
the 20 employees of your small business. You might chart the employees'
ages to get a better grasp on whether there are any large clumps of
employees who might retire at similar times. Or if you have an employee
you suspect is under-performing on the job, you might randomly select 15
other employees and chart their productivity (in terms of customers served)
with that of the suspected employee to see if his falls far below or in line with the others'. Both of
these exercises would give you a good sense of what the data is telling you.

A dot plot shows each data value as a point, distributed along a horizontal axis. Dot plots are useful
because they show the distribution of a data set, as every data value is represented by a dot.

Example
The table below shows the number of computer lab visits each day for20 days. Construct a dot
plot for the data.

43 51 13 31 20
32 32 23 32 2
52 57 44 36 33
44 32 45 14 25

Step 1

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Arrange the data set in order from least to greatest.

2, 13, 14, 20, 23, 25, 31, 32, 32, 32, 32, 33, 36, 43, 44, 44, 45, 51, 52, 57

Step 2

Draw a horizontal axis with a scale from0 to 60.

Step 3

Place a dot above the horizontal axis for each data point in the table. Here, each "dot" is
represented by the letter x. Any repeated values, (such as 32 or 44 which have red boxes around
them), should be represented by a mark for each value, stacked vertically.

Stem plots
Stem plots, also called stem-and-leaf plots, are another way to show a data set and its distribution
or shape. A stem plot is constructed by separating each data value into a stem (usually the left-most
digit) and a leaf (usually the right-most digit). For example, if 36 is a data point, the 3 would be the
stem and the 6 would be the leaf. The data is arranged in two columns, stems and leaves, with a
vertical line separating the columns.

Example
Using the same data from above, a stem-and-leaf plot would look like this:

Stems Leaves
0 2
1 34
2 035
3 1222236
4 3445
5 127

Exercise

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Dot Plots
The following dot plot, which shows the results of a recent network test of download
speeds, will be used for Questions 1-5:

1. Refer to the dot plot above. What is the minimum value in the data
set?
2. Refer to the dot plot above. What is the maximum value in this
data set?
3. Refer to the dot plot above. Is the value of 0 in the data set?
(Enter Yes or No)
4. Refer to the dot plot above. During this network test how many
reports recorded a download speed of 60 Mbps?
5. Refer to the dot plot above. What download speed was most often
recorded during this network test?
Stem Plots
The following stem plot will be used for Questions 6-10:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


6. Refer to the stem plot above. What is the minimum value in the
data set?
7. Refer to the stem plot above. What is the maximum value in this
data set?
8. Refer to the stem plot above. What is the most frequent age
represented in this data set?
9. Refer to the stem plot above. How many stems are represented in
this stem plot?
10. Refer to the stem plot above. Is the value 34 contained in this
data set? (Enter Yes or No)

4.04.2 Histograms
Histograms

Career Connections

Histograms for Class Distributions

Teachers have very diverse classrooms. One of the exciting challenges


teachers face is how to teach to the needs of different types of students. A
histogram can help with teaching to these needs.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


A histogram is a tool for describing the distribution of a single quantitative variable, like test scores.
Teachers and administrators use histograms to see how students' test
scores are distributed. For example, if a histogram for Mr. Smith's classroom
shows a lot of students scoring in the lower range, he might make the case
to his principal that a teacher's aide is needed to provide more individual
instruction to struggling students. On the other hand, a histogram for Ms.
Bellamy's science class might show a lot of students scoring in the higher
range. This might help her understand how she could target more advanced
topics for her students, who are already performing very well.

In this lesson, you'll learn about working with histograms from many different disciplines, such as
business and healthcare. In the end, histograms can be used to help us understand a single
quantitative variable better. Be sure to think about how you could use histograms in your own
discipline as you go through this material.

A histogram is a graph that displays quantitative data. The vertical bars in a histogram show the
counts or numbers in each interval. A comparison of the intervals, or a review of the graph as a
whole, helps the audience understand the information presented.

The distinction between a histogram and a bar chart is an important distinction to make. As
previously discussed a bar chart measures categorical data that is distributed over groups or
categories, while a histogram measures how quantitative data is distributed over various intervals.
For example, a histogram would be appropriate to display how many people fall in various intervals
of heights, as height is an example of quantitative data.

In other words, a histogram is used to display frequencies or relative frequencies for quantitative
data; in contrast, a bar chart is used to display frequencies (i.e., counts) or relative frequencies for
categorical data.

Histograms vs. Bar Charts

Histogram: displays frequencies or relative frequencies for quantitative data

Bar Chart: displays frequencies (i.e., counts) or relative frequencies for categorical data

Histograms allow team members and stakeholders to view a significant amount of data at one time,
and to see how data is distributed across various intervals of values. The histogram's bars
represent the values or intervals in the study. The height of each bar shows how many
observations or events fall into each interval. The shape of the graph illustrates how the data is
distributed.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


To create a histogram, data is collected using a check sheet. It is necessary to decide what
intervals, or values, you are going to use and note how the data is distributed.

The intervals in a histogram need to encompass all of the data collected. It is important to make
sure the minimum and maximum values are accounted for, as well as every value in between.
Additionally, make sure that the intervals are comparable and exhaustive. The intervals should run
consecutively so that all data is accounted for and the visual representation of the graph is
accurate.

It is also important to use an appropriate number of intervals; if you have difficulty determining how
many intervals to use, you can refer to the rough guidelines laid out in the chart below:

If you have: Divide the data into:


Fewer than 50 measurements 5 to 7 intervals
50 to 100 measurements 6 to 10 intervals
100 to 250 measurements 7 to 12 intervals
Greater than 250
10 to 20 intervals
measurements

Data Distributions: Symmetry versus Skewness

Distributions are often categorized into two different types.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Symmetric Distributions

A symmetric distribution is a common type of frequency distribution. As the name indicates,


symmetric distributions are symmetrical, with the left half of the histogram being roughly
identical to the right half. In other words, if you cut the histogram in half, each side would
be a near-perfect mirror image of each other.

This symmetry, or type of distribution, is illustrated in the histogram below. Notice how the
middle value(s) is the most frequent. The values decrease in a symmetrical manner, to the
right and left of the center of the histogram.

A bell curve, or normal distribution has a very specific distribution. Normal distributions
will be discussed later in this module. For now, it is important to note that just because a
histogram is symmetric does not make it normal!

Skewed Distribution

Skewed distribution

Distributions can also be asymmetric. A skewed distribution is the term used to describe a
distribution that has a "long tail" on one side of the peak. In other words, the distribution is
lopsided and does not have a symmetrical shape. Skewness is used to measure the
asymmetry of a distribution. When more data falls further to the left of the peak, it is known
as skewed left. This type of distribution is also referred to as negatively skewed.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Similarly, a data set can be skewed right, when more data falls further to the right of the
peak of the histogram. This type of distribution is also referred to as positively skewed.

A histogram that is skewed right, or positively skewed indicates that more values are far
greater than the most common value, but not far less. A histogram that is skewed left, or
negatively skewed, indicates that more values are far less than the most common value,
but not far greater.

Career Connections

Histograms

As with other graphs, histograms and bar charts help people grasp and interpret data. Therefore,
the charts can be used to organize data to see if recognizable patterns emerge. For example, if you

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


are working at a clothing company and considering how to price next
season's line, you might be interested to learn the price points at which last
season's merchandise sold.

When you plot the number of items sold at a certain price points on a histogram, you might notice
that they cluster at $ 20 to $ 29 and $ 80 to $ 89. More expensive items, originally priced at $ 40 to
$ 50 did not sell until they were discounted into the$ 20 to $ 30 range while merchandise originally
priced at $ 80 to $ 90 (as well as more expensive merchandise discounted to that range) sold
relatively well. Using this information, you might design or price more items to fall into the popular
ranges, hoping to capitalize on your customers' comfort level with these price points.

While a tabular record of this data would convey the same information if carefully inspected, the
height of the bars at $ 20 to $ 29 and $ 80 to $ 89 leaps out at the viewer. The relative differences of
each bar can be more easily grasped than a set of numeric data.

4.04.3 Describing the Distribution of Quantitative Data


Describing the Distribution of Quantitative Data
Dot plots, stem plots, box plots, and histograms are all effective visual displays of quantitative data.
Along with visually displaying quantitative data, it is also important to be able to interpret these
graphs, and describe the quantitative data. In describing a distribution of data illustrated in a graph,

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


it is important to note:

Graphical Displays: Describing Distributions

The shape of the graphical display


The spread of the data
The maximum and minimum values
Any values that could be outliers

Different distributions hold different shapes. As we have seen earlier a distribution can be
symmetric:

A distribution can be skewed-right:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


A distribution can be skewed-left:

There are also other common shapes distributions can hold. As well as bell-shaped distributions,
another common symmetric distribution is U-shaped. A U-shaped distribution occurs when a
symmetric distribution has a "valley" rather than a peak:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Another common symmetric distribution is known as uniform. A uniform distribution has similar,
uniform frequency across its range:

If a distribution has two clear peaks rather than one, it is known as bimodal. A histogram with two or
more clear peaks is called multimodal. Below is an example of a histogram that is bimodal.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Being able to describe the shape of a distribution is very important. This terminology provides a
common language with which we can discuss the appearance of quantitative data.

Importance of Center and Spread

When describing quantitative data, center and spread are two important characteristics. The center
of a set of quantitative data is a point that represents the "middle" of the data. As we will see, this
can be measured in many different ways. There are also multiple measurements that are used to
describe spread. Spread, in general, is a way to describe the dispersion of quantitative data. Is all of
the data clustered around one point, or is it spread out?

Consider the following data set:

We can create a graphical display in order to better understand the shape, center, and spread of the
data. A histogram would be a good choice.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


This histogram gives us a good sense of the shape of the data, including the approximate center
and spread of the data. We can also see that the data is skewed to the right.

Career Connections

The Importance of Recognizing Normal versus Skewed Distributions

It is important to pay attention to the skew of the distribution if you are


planning to use the mean (the average) in a business context. The mean is
not "really" in the middle of a data set if there are trailing points to the far left
or far right because these trailers wield an outsized influence on the mean.

For example, suppose you measure your employees' average salary to


evaluate whether you are paying a living wage for your region. You may be
pleased to discover that employees' average salary is just under $ 50, 000, a good wage for the
region your company is based in. However, if the CEO is paid 20 times the amount paid to the least
paid employee, that salary alone may skew the entire data set to the right (higher). Once you
eliminate the CEO's salary, you may discover that the mean salary paid is closer to $ 35, 000. In a
larger company, the same effect can happen if you have a highly compensated C-suite. The large
salaries for a very few people can raise the average salary for the company as a whole in a way
that misleads observers into thinking the company is paying better wages (on average) than it
really is.

Exercise

Exercises: Distribution
Answer the question below by analyzing the following histogram. Enter the letter that corresponds
with your answer.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


1. The histogram represents data that is:

A. Skewed left

B. Skewed right

C. Uniformly distributed

D. Normally distributed
2. Data that is skewed right is:

A. Negatively skewed

B. Positively skewed

C. Symmetric

D. Normally distributed
3. Which of the following does not describe a symmetric histogram?

A. Bell-shaped

B. U-shaped

C. Positively Skewed

D. Uniform
Answer the questions 4 and 5 below by analyzing the following histogram. Enter the letter that
corresponds with your answer.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


4. The histogram above is known as which of the following?

A. Bimodal

B. Skewed left

C. Uniform

D. Skewed right
5. What can we infer about this data?

A. There are two peaks in the data set

B. The data is negatively skewed

C. The data is positively skewed

D. All intervals in the distribution have the same number of


observations
Answer the question below by analyzing the following histogram. Enter the letter that corresponds
with your answer.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


6. How would you describe the distribution above:

A. Bimodal

B. Skewed left

C. Skewed right

D. Symmetric
7. A histogram with a symmetric distribution that has a "valley" rather
than a peak is described as:

A. Unimodal

B. U-shaped

C. Uniform

D. Bell shaped

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Answer the questions below by analyzing the histogram. Enter the letter that corresponds with your
answer.

8. This histogram represents data that is:

A. Skewed left, negatively skewed

B. Skewed left, positively skewed

C. Skewed right, negatively skewed

D. Skewed right, positively skewed

9. Data that is skewed left is:

A. Negatively skewed

B. Positively skewed

C. Symmetric

D. Normally distributed
10. A histogram that has more than two modes is known as:

A. Unimodal

B. Uniformly distributed

C. Multimodal

D. Symmetric

4.05 Numerical Measures


Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.
4.05 Numerical Measures
Numerical Measures
As we review the basics of mathematics, we find ourselves at one of the most useful branches of
math: statistics. Statistics provides the tools that are used to analyze and summarize large
quantities of numerical data. With statistics, we can interpret and present data in a concise,
comprehensive, and easily understandable way.

A data set is any collection of numerical values, such as measurements, observations, or survey
responses. For example, if we measure the heights (in centimeters) of ten randomly selected
people, we could have the following data set:

Reliability and Validity of Data

In statistics, measurements need to be both reliable and valid. Reliable data is both consistent and
repeatable. If you were to administer the same test to the same person three times and the scores
were similar each time, the test could be categorized as reliable. If the results varied greatly, the
test would be unreliable. Similarly, valid data is data resulting from a test that accurately measures
what it is intended to measure. For instance, if a test reflects an accurate measurement of a
student's abilities, it is said to be valid.

Communicating Calculations

Previously, we've roughly estimated the center, spread, and possible outliers of data. For more
precise results, we calculate these measures.

Reliable and valid data calculations give rise to more accurate and precise results. This precision
allows for additional graphical displays for quantitative data. Rather than rough estimations, we can
display and report precise figures that measure spread, center, and other summary values of data.

4.05.1 Measures of Center

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Measures of Center
A numerical summary is a number used to describe a specific characteristic about a data set.
Numerical summaries can be used to define a measure of central tendency, or the measure of the
center of a data set. Measures of the center of a data set are the mean, median, and mode.

Mean

The mean is one of the most useful measures of central tendency. The mean, also known
as the average, is a single value that represents the center of a set of data values. Mean
can be substantially influenced by one or more extreme values in a data set (think skewed
data), so mean is only used when the data is symmetric. Therefore, we say that the mean
is not a resistant measure of center.

To calculate the mean, the values in a data set are simply added together and divided by
the number of available values.

Let's use the ten heights from the example on the previous page.

Step 1

To find the mean, the first step is to add all of the values together.

172. 7 + 168. 3 + 182. 9 + 167. 6 + 189. 2 + 177. 8 +

185. 4 + 166. 4 + 193. 7 + 165. 1 = 1769. 1

Step 2

Next, you divide that sum (1769.1) by the number of values in the data set 10).
(

1769. 1 ÷ 10 = 176. 91

The mean of the above data set equals 176.91.

Example
Mr. Nolan coaches a little league baseball team made up of 18 players. On the team, there
are 48-year-olds, 27-year-olds, 36-year-olds, 39-year-olds, 210-year-olds, 35-year-olds, and
112-year-old. Eleven of the players are boys and seven are girls. What is the mean age on
Mr. Nolan's baseball team?

To calculate the mean, sort out the data that's applicable to the team ages. Though within
the problem those of the same age are grouped together, to find the mean, we need to
account for each of the 18 players' ages individually. The gender of the players is not
relevant in this problem.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Now that we have the entire data set, let's add up all of the ages on the team.

8 + 8 + 8 + 8 + 7 + 7 + 6 + 6 + 6 + 9 + 9 + 9 + 10 + 10 +
5 + 5 + 5 + 12 = 138

Next, we divide 138


(the sum of the players' ages) by 18
(how many players are on the team).

138 ÷ 18 = 7. 6666. . .

When we round to the tenths place, our answer is 7. 7


. So our mean (or average) age on Mr. Nolan's baseball team is
7. 7 years old.

Median

The second measure of central tendency is the median. The median is the "halfway" point
of a set of values; an equal number of values will fall above and below the median of a
data set.

Unlike the mean, the median is not overly influenced by extreme values in the data set, so
we can use the median when the data is skewed. Therefore, we say that the median is a
resistant measure of center. To properly find the median, values must be first sorted from
smallest to largest.

Odd Number of Values in a Data Set

Below is a set of values representing business credit scores.

Step 1

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


First, sort the data from lowest value to greatest value in order to find the median.

Step 2

Next, count the total number of values in the data set to find the median. If there is anodd
number of values, the halfway point will fall directly on a value and that will be your
median.

There are 15
values in total, which means the median, or halfway point of the data set, is 72
. (There are seven values below 72
, and seven values above 72
.)

Even Number of Values in a Data Set

Now, let's say we have a data set with only14


values.

In this data set, the median would fall in between72


and 73
(indicated by the red line). In the case of an even number of total values, the median is
the halfway point between the two middle values.

This median can be calculated by adding those two middle values and dividing by two.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


(72 + 73) ÷ 2

= (145) ÷ 2

= 72. 5

So the median of the second data set is72. 5

Example: Student Logins


A data set for daily student logins to an online course is shown below:

Step 1

Sort your data from smallest to largest.

Step 2

Imagine drawing a line down the middle of your data so that half the data points are on the
right of the line and half are on the left.

If your line lands on a data point (for example, if you have an odd number of values), that
is the median.

If your line lands in between two data points (for example, if you have an even number of
values), your median is the halfway point between the two data points, and the median is
calculated by taking the average of those numbers. In this case:

(50 + 51) ÷ 2 = 50. 5

Mode

The mode is the third and final measurement of central tendency. The mode represents the
value that occurs most often in a data set.

The mode is only relevant if a data set has values that are repeated, and unlike mean and
median, there can be more than one mode in a data set.

Below is a set of values.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


When solving for the mode, all that matters is which value appears most often. In this case,
the mode would be 30
.

Example
Mode is a measurement easily displayed in graphs -
also unlike mean and median. Take a look at the histogram below, which presents a range
of test scores, and decipher which bar in the graph displays the mode.

Notice how the bar above Score of 90 %


is the tallest. With the given data, we can assume 12
students scored a 90 %
on the test—more than any other test score achieved. Therefore, the mode of this data
would be Scores of
90 % because that value appears most often.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Limitations of Mode
The mode has the following important limitation: There can be more than one!

1. In fact, if all of the intervals contained the same number of scores, every value would
be the mode, rendering that measure useless. For this reason, you should always
graph your data. The other two measures of central tendency, median and mean,
may not tell you this information about a data set's distribution as well as a graph.

Below are three graphs, each with a mean of approximately 3. 33


. Notice the difference in shape. The graph on the left is unimodal, the graph in the
middle has a uniform distribution, and the graph on the right is bimodal.

Career Connections

Mean, Median, and Mode

Those working in management are often asked to review research to help


determine the current policies and procedures in order to come up with
improved systems. Understanding the mean, median, and mode is
important, as they are some of the basic measurements used in research.

Employees working in management conduct performance improvement


(PI) or continuous quality improvement (CQI). Both of these activities look
at processes or activities to help identify ways of improving them.
Collecting and reviewing data are key steps in both of these activities. An example of this would be
an industrial unit may look at the time it takes a worker to complete the task of creating a
component. Each step in the process will be timed and entered into a database. Data is collected
for a preset period, such as 30
days. Once all of the data is collected, one of the first steps in the data analysis would be to
determine the mean, median, and mode. The average time of each step would be the mean; the
median is the time that appears halfway in the data set, and the mode is the most common time in

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


the data set.

4.05.2 Measures of Spread


Measures of Spread

Career Connections

Range and Spread

It is always important to pay attention to the spread of the data. Suppose


you have just been hired to manage a dog grooming business. Your first
week on the job, you gather data about how long it takes to complete
various routine tasks. For dog washing, you find a wide spread in the task
from a minimum of 5
minutes to a maximum of 45
minutes. Even the same employee can take very different amounts of time
to wash different dogs. On the other hand, toenail clipping clusters strongly around 10
minutes, with a minimum of 8. 5
minutes and a maximum of 11
minutes.

What do these two data sets tell you? First, it is reasonable to expect that any employee should be
able to clip toenails in 11
minutes. When you create the employee schedules, you know pretty precisely how much time to
allow. Dog washing, however, requires further investigation. What makes the difference in the
times? In a new data set, you might chart the time taken to wash the dog against its weight (a
proxy for size) or against its manageability. When you have identified the cause of the spread, you
can ask for the relevant information from the customer and use it to create efficient schedules for
your dog groomers. Until you have a better grasp of how to predict the amount of time needed for
dog washing, you will have difficulty creating a schedule.

In addition to determining the center of a distribution by using a measure of central tendency, it is


also helpful to know how spread out the data are, or how variable the data are. Measures of
spread (variability), including the range and the standard deviation, can tell you this. As previously
discussed, the mean, median, and mode describe the center of a data set. Similarly, we use the
range and the standard deviation to describe the spread of a data set.

Range

Range

Range is the difference between the smallest (minimum) and greatest (maximum) values of
a data set.

The minimum is the smallest value available in a data set. The maximum is just the
opposite; it's the greatest value in a data set.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Minimum, maximum, and range are measurements often used to bring clarity and scope to
a set of values.

Example

Below is a set of data.

18, 11, 3, 26, 13, 40, 31, 5, 12, 45, 52, 22, 17, 33, 8

To find the minimum, maximum, and range, it is helpful to first sort the values from smallest
to largest.

3, 5, 8, 11, 12, 13, 17, 18, 22, 26, 31, 33, 40, 45, 52

After the data is sorted, the minimum and maximum fall at either end of the data set. In the
data set shown, the minimum is 3
and the maximum is 52
.

To find the range, simply subtract the minimum from the maximum.

Range =
maximum -
minimum

52 - 3 = 49

The range of the above data set equals 49


.

Interquartile Range

Quartiles are widely used measures when dealing with data sets. Quartiles are values that
divide a data set into four equally sized groups. There is one median per dataset that splits
the data into two equally sized groups. Similarly, a dataset has three quartiles that split the
data into four equally sized groups. The interquartile range measures the difference
between the third quartile and the first quartile. To illustrate how to find the first and third
quartiles we will use the data from the following research poll that asked individuals how
many alcoholic drinks they have per week.

Example
Polling 12
people about how many alcoholic drinks they have per week might yield the following
data:

9, 1, 7, 5, 4, 3, 6, 2, 9, 1, 9, 4

First, from least to greatest:

1, 1, 2, 3, 4, 4, 5, 6, 7, 9, 9, 9

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


To determine the first and third quartiles, order the data from lowest value to highest value.
Then separate the data into four equal groups.

As we have 12
values in our data set, placing the data into the four quarters results in having three data
values in each quarter. As you can see in the chart above, specific numbers that have
multiple occurrences are included the number of times they occur.

The interquartile range is an indicator of the distribution of a sample and can also help
identify any outliers. Outliers are data points (numbers) that are far away from all other data
points. It is helpful to identify any outliers and determine whether they should be used.

Follow these steps to find the interquartile range of a data set.

Steps to Find the Interquartile Range

1. Put the data set in order from least to greatest.

1, 1, 2, 3, 4, 4, 5, 6, 7, 9, 9, 9

2. Find the median, or midpoint, of the data set. This can also be called the second
quartile (Q2
).

1, 1, 2, 3, 4, 4
| 5, 6, 7, 9, 9, 9

Q2 = 4. 5

3. Identify the median of the lower half of the data set and label it asQ1
(the first quartile).

1, 1, 2

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


| 3, 4, 4
| 5, 6, 7, 9, 9, 9

In this case, the median of the lower half of the data set is midway between2
and 3
, which averages to 2. 5
.

Q1 = 2. 5

4. Identify the median of the upper half of the data set and label it asQ3
(the third quartile).

1, 1, 2
| 3, 4, 4
| 5, 6, 7
| 9, 9, 9

In this case, the median of the upper half of the data set is midway between7
and 9
, which averages to 8
.

Q3 = 8

5. Subtract Q1
from Q3
to determine the interquartile range, or IQR.

1, 1, 2,
| 3, 4, 4,
| 5, 6, 7,
| 9, 9, 9

IQR = Q3
-
Q1

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Q3 = 8

Q1 = 2. 5

IQR
= 8 - 2. 5 = 5. 5

IQR
= 5. 5

Let's look at one more example of finding the IQR for a dataset.

Example
Find the IQR of the following set of download speeds from various clients of a certain
ISP.

21, 8, 41, 20, 11, 16, 13, 14, 35, 27, 40, 18, 30, 32, 20

Follow these steps to find the interquartile range of a data set with an odd number of data
points.

Steps to Find the Interquartile Range — odd number of data points

1. Put the data set in order from least to greatest.


8, 11, 13, 14, 16, 18, 20, 20, 21, 27, 30, 32, 35, 40, 41

2. Find the median (Q2


) of the data set. Since there are15
data points, the median will be the middle value, which is the 8th
value in the data set.

8, 11, 13, 14, 16, 18, 20, 20, 21, 27, 30, 32, 35, 40, 41

Q2 = 20

3. Identify the median of the lower half of the data set and label it asQ1
(the first quartile). In this case, there are 7
data points in the lower half of the data, so Q1
will be the 4th
in that list:

8, 11, 13, 14, 16, 18, 20, 20, 21, 27, 30, 32, 35, 40, 41

Q1 = 14

4. Identify the median of the upper half of the data set and label it asQ3
(the third quartile). In this case, there are 7
data points in the upper half of the data, so Q3

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


will be the 4th
in that list:

8, 11, 13, 14, 16, 18, 20, 20, 21, 27, 30, 32, 35, 40, 41

Q3 = 32

5. Subtract Q1
from Q3
to determine the interquartile range, or IQR.

IQR = Q3
-
Q1

Q3 = 32

Q1 = 14

IQR
= 32 - 14 = 18

IQR
= 18

The 1.5 IQR Criterion: Identifying Outliers

Outliers are defined to be any points that are more than1. 5


×
IQR
above Q3
or below Q1
. This rule that is used to identify outliers in a data set is called the1.5 IQR Criterion Rule
. The following graphic illustrates how outliers fall either to the left ofQ1
-
1. 5 IQR
, or to the right ofQ3
+
1. 5 IQR
.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Building on the previous example, let's see if there are any outliers in this data set.

Steps to Identify Outliers in a Data set

1. Recall the data set from the previous example, placed in order from least to
greatest.

1, 1, 2, 3, 4, 4, 5, 6, 7, 9, 9, 9

2. The interquartile range, or IQR


, for this data set is:

Q3 = 8

Q1 = 2. 5

IQR =
Q3
-
Q1 = 8 - 2. 5 = 5. 5

IQR = 5. 5

3. Multiply the IQR


by 1. 5
.

IQR × 1. 5 =

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


5. 5 × 1. 5 = 8. 25

4. Add the result to Q3


:

Q3 = 8

8 + 8. 25 = 16. 25

Any values greater than this number,


16. 25, are outliers.

5. Subtract the result from Q1


:

Q1 = 2. 5

2. 5 - 8. 25 = - 5. 75

Any values less than this number,


-5. 75, are outliers.

6. Review the data set:

1, 1, 2, 3, 4, 4, 5, 6, 7, 9, 9, 9

Given there are no data points less than


-5. 75, or greater than
16. 25, this data set does not contain any outliers.

Five-number Summary

The five-number summary lists the minimum, first quartile, median, third quartile, and
maximum in a data set. The five number summary will be represented by a graphical
display that we will learn about later in Module 4 called a box plot.

The following table is a five-number summary of the hours spent volunteering in the past
year taken from a sample population of employees.

Five-number Summary
Statistic Data Value

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Minimum 105
Q1 119. 5
Q2
126. 5
(Median)
Q3 138. 5
Maximum 169
Standard Deviation

The standard deviation tells you how far, on average, the data points are from the mean. In
this course, we will not focus on how to calculate the standard deviation (we can use
computers to do this for us) but rather on building an intuition for using the standard
deviation to measure how spread out the data is in a dataset. Standard deviation is a
measurement that is used for symmetric data.

As displayed in the graph below, in a bell-shaped curve, also known as a normal


distribution, the Standard Deviation Rule states that 68 %
percent of the data will fall within 1
standard deviation of the mean, 95 %
of the data will fall within 2
standard deviations of the mean, and 99. 7 %
of the data will fall within 3
standard deviations of the mean. A greater standard deviation means that the data is more
spread out.

The Standard Deviation (Empirical) Rule

Approximately 68 %
of all values are within 1
standard deviation of the mean
Approximately 95 %
of all values are within 2

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


standard deviations of the mean
Approximately 99. 7 %
of all values are within 3
standard deviations of the mean

From the Standard Deviation Rule, we can calculate all of the parts of the bell-curve. It is
important to memorize the Standard Deviation Rule. You should also know how to
calculate the other percentages, or you can memorize those.

Let's take a look at a couple examples to see how we can apply the Standard Deviation
Rule.

Example
Custom log home builder, Tinker Log Homes Inc. takes an average length of40
weeks, or 280
days, with a standard deviation of 13
days to build a custom log home (assume a normal distribution). Knowing this information,
and using the Standard Deviation Rule, let's answer the following questions.

1.
68 % of the data is between what two values?

The curve above illustrates this distribution. Using the Standard Deviation Rule, we
know that 34 %
of the data is one standard deviation above the mean (+1
SD), and 34 %
of the data is one standard deviation below the mean (-1
SD). Therefore, according to the Standard Deviation Rule, 68 %

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


of the data falls between one standard deviation from the mean.

Using this information, we can now calculate the values that68 %


of the data will fall between:

Answer: 68 %
of the values will fall between 267
and 293
days.

2. What percentage of custom log home build timelines will range between 254
and 306
days?

To answer this question, first, we have to figure out how far both254
and 306
are from the mean.

The difference between both of these values and the mean is equal to26
. Next, we divide this value by our standard deviation of13
. The result is equal to 2
. Therefore, 254
and 306

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


fall two standard deviations away from our mean of 280
.

Now that we know that these two values are two standard deviations from the mean
of 280
days, we can use the Standard Deviation Rule to determine what percentage of
home build timelines will last between 254
and 306
days.

As illustrated in the distribution curve above, the Standard Deviation Rule states that
13. 5 %
of the data falls between one and two standard deviations from the mean. Knowing
this information, we can now calculate the percentage of data that falls between two
standard deviations (or between 254
and 306
days) from the mean by adding up the percentage values in the shaded areas of the
curve.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Answer:
95 % of the values will fall between254
and 306
days. In other words, the Standard Deviation Rule tells us that95 %
of all data will fall between two standard deviations from the mean.

3. What is the chance that a custom log home build will last longer than306
days?

Since we know that 100 %


of the data is included in the bell curve and that the curve is symmetric, we can also
find the percentages for parts of the curve. From the previous question, we know
that 306
is two standard deviations from the mean. To answer, we want to calculate the
percentages of the parts of the curve that are greater than two standard deviations
from the mean.

The percentage values under the curve that are greater than two standard
deviations from the mean are illustrated in the distribution curve below.

Next, we add together the two percentage values:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Answer: 2. 5 %
of custom log home builds will last longer than 306
days.

4.05.3 Outliers and Choosing Measures


Outliers and Choosing Measures
When looking at data, it is productive to look for outliers—observation points (numbers) that are
distant from other observations. When outliers are detected, we can try to figure out what leads to
these results. We examine our collection techniques, and the data itself, and determine whether the
number is incorrect or faulty in some way. We can also determine whether a figure is an outlier
because it describes something that does not belong in the study. An outlier may even be correct,
but it is helpful to identify any outliers and determine whether they should be used, or whether their
existence is indicative of something that has gone wrong.

Extreme Values in a Symmetric Distribution

If a distribution is symmetric, such as a bell-curve, a u-curve, or a uniform distribution, there will be


roughly similar extreme values on either side of the distribution. That is to say that a symmetric
distribution will have similarly extreme positive and negative values. Examine the bell-curve below:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Here, we know that there is equal "pull" coming from the data on both sides. As a feature of
symmetric distribution, there should not be many extreme outliers on one side of the distribution but
not the other. Therefore, our measures of center (mean, median, mode), will all be roughly in the
middle of the histogram, at the peak of the data. Skewed data is often more complicated to
measure than symmetric data, though.

Extreme Values in a Skewed Distribution

We refer to values in a histogram that come after a gap as extreme values and possible outliers.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Even if the data is skewed, that does not necessarily mean that the data to the far right (or left) are
outliers. What is important to note is how outliers influence measures of center and spread.

Examine a skewed-right distribution below:

Due to the fact that this distribution skews to the right, the most extreme values are on the right side
of the distribution. These extreme values, or possible outliers, have an effect on the measures of
center.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


The mode remains at the peak of the distribution. The median is always in the middle of the data.
The mean is pulled to the right, as the most extreme values are far right of the center of the
distribution, causing them to have a dramatic effect on the mean.

Imagine you are looking at income data among a group of people. Most of the people in this data
set are making between $30, 000
and $75, 000
per year. There are a few high-earners who have incomes over $100, 000
and there is one person who is earning $10, 000, 000/
year. This distribution is skewed-right. The few individuals with a very high income skew the mean
to the right. You may find that the mode income is $40, 000
, the median income is $55, 000
, but the mean income is $100, 000
. In this sample, $40, 000
is the most common income, and the middle-earner makes $55, 000/
year. The mean income is less useful, though, as the highest earners have dramatically skewed the
mean to the right.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


When the data skews in one direction or another, it is often best to use median and interquartile
range to describe the center and spread of the data. Mean and standard deviation are more
appropriate measures of center and spread for symmetrical data.

The following table summarizes the preferred measures of center and the measures of spread for
normal and skewed distributions.

Measures of Measures of
Distribution
Center Spread
Skewed Median Range or IQR
Normal Symmetric Mean Standard Deviation

4.06 Review: Measures of Center and Spread

4.07 Additional Graphical Display for Quantitative Data: Box


Plots
Additional Graphical Display for Quantitative Data: Box Plots
Box plots, also known as box-and-whisker plots or hinge plots, show the distribution or shape of a
data set. Box plots are useful when showing the variability and spread within a data set because
they take into account median and quartiles rather than averages. (Recall the discussion on
quartiles on assignment 4.05.2)

Anatomy of a Box Plot

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Box plots have four parts: the first whisker, two rectangles, and another whisker. Regardless of size,
each part represents 25 %
of the data.

A modified box plot is just like a regular box plot, except that outliers are shown as points above or
below the minimum and maximum. The box plot below illustrates an outlier. In this example, the
outlier is beyond the end of the whisker.

Five-number Summary

A box plot is a convenient way to show five important statistical values: minimum, maximum, first
quartile, median, and third quartile. As previously mentioned, these five values are often referred to
as a five-number summary of a data set. Many statistical outputs from technology, such as
calculators, computer programs, or other software, have five-number summary displays built in.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Box plots can either be displayed horizontally or vertically as illustrated in the following two
examples.

Example: Horizontal Box Plot


Consider the following box plot that illustrates the distribution of Algebra exam scores for cohort A.

Five-number Summary

The far left side of the box plots represents the minimum value of the data set. In the
above example, this is 60
.
The left whisker represents the lower 25 %
of the data. In this example, that is from 60
to 70
. That means 25 %
of students scored between 60
and 70
on the algebra exam.
The first part of the box represents the next 25 %
of the data. It is shaded blue in the above box plot. So another 25 %
of the students scored between the Q1
of 70
and the median of 75
.
The line in the middle of the box is the median Q(2
). That means 50 %
of the data is below this point, and 50 %
of the data is above this point. In the above example, 50 %

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


of students scored below 75
and 50 %
scored above 75
.
The second part of the box represents the next 25 %
of the data. It is shaded orange in the above box plot. In our example, we see that 25 %
of the students scored between the median of 75
and the Q3
of 85
.
The right whisker represents the upper 25 %
of the data. In the above example, we can see that 25 %
of students scored between 85
and 100
on the algebra exam.
The far right side of the box plots represents the maximum value of the data set. In the
above example, this is 100
.

Example: Vertical Box Plot


Consider the following vertical box plot that illustrates the lengths found in a population of sharks.

Five-number Summary

The top of the line is the maximum value in the data set; in the above example, the
Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.
The top of the line is the maximum value in the data set; in the above example, the
maximum value is approximately 18
.
The upper limit of the rectangular box, shaded in purple in the above example, represents
the third quartile (Q3
).
The line in the middle of the rectangle (in the above example, this line is indicated where
the green and purple shaded areas of the rectangle meet) is the second quartile (Q2
) or median.
The lower limit of the rectangular box (in the above example, shaded in green) represents
the first quartile (Q1
).
The bottom of the line represents the minimum value of the data set; in the above
example, this is approximately 7. 5
.

Exercise

Box Plots
The box plot below shows the distribution of the time a real estate agent spends with a client prior
to the client closing on a new home. Answer the questions below by analyzing the box plot.

1. What is the median number of days an agent spends working with


a client?
2. What is the minimum value in the data set?
3. What is the maximum value in this data set?
4. What is the value of the first quartile?
5. What is the value of the third quartile?
6. Below is a data set of the total sales during the past quarter recorded by the20
salespeople in the Northeast Region. To complete the exercise, fill in the blanks with the
corresponding answer. Enter only numerical digit(s) such as "5
" in the blank.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Calculate the five-number summary of this data: minimum, first quartile (Q1
), second quartile (median), third quartile (Q3
), and maximum.

Minimum:
Q1
:
Median:
Q3
:
Maximum:
7.

Identify the Five-number Summary values by entering in the letter on the graph that corresponds to
that value.
Minimum Value:
Q3
:
Maximum Value:
Q1

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


:
Q2
:

Review Checkpoint
To test your understanding of the content presented in this assignment, please click on the
Question icon below. Click your selected response to see feedback displayed below it. If you have
trouble answering, you are always free to return to this or any assignment to re-read the material.

1. Which of the following represent the five values included in the five-number summary of a data
set?

a. minimum, first quartile, mean, third quartile, maximum

Incorrect. Try again.

b. minimum, first quartile, median, third quartile, maximum

Correct.

c. minimum, mode, mean, median, maximum

Incorrect. Try again.

d. first quartile, mode, mean, median, third quartile

Incorrect. Try again.

2. Based on the following box plot, what percent of initial client consultations last less than70
minutes?

a. 75 %

Correct. The answer is a. We know that each part of the box plot represents25 %
of the data and that 70

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


minutes represents Q3
. As 75 %
of the data falls below Q3
, we can determine that 75 %
of initial client consultations last less than 70
minutes.

b. 50 %

Incorrect. Try again.

c. 40 %

Incorrect. Try again.

c. 25 %

Incorrect. Try again.

3. Based on the box plot below, approximately what percent of salespeople represented in this data
set have a total first quarter sales level less than 120
?

a. 10 %

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Incorrect. Try again.

b. 15 %

Incorrect. Try again.

c. 25 %

Correct. Looking at the box plot, Q1


is approximately 120
. As we know 25 %
of the data falls below Q1
, we can determine that 25 %
of the salespeople represented in this data set will have a total first quarter sales level below 120
.

d. Cannot be determined

Incorrect. Try again.

4. Based on the box plot below, what percent of the class scored between75
and 85
on the algebra exam?

a. 25 %

Correct. From this box plot we see the median value is equal to75
, and Q3
is equal to 85
. Since we know in a box plot that25 %
of the data falls between the median and Q3
, we can determine that 25 %
of the class scored between 75

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


and 85
on the algebra exam.

b. 10 %

Incorrect. Try again.

c. 50 %

Incorrect. Try again.

d. 15 %

Incorrect. Try again.

5. Based on the box plot below, what percent of the data falls between83
and 88
?

a. 5 %

Incorrect. Try again.

b. 25 %

Incorrect. Try again.

c. 35 %

Incorrect. Try again.

b. 50 %

Correct. We know that 50 %


of the data represented in a box plot falls between the median and the maximum value. In this box
plot, 83
is the median and 88

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


is the maximum value. Therefore, 50 %
of the data values fall between these two values.

4.08 Exercise: Identifying Graphical Displays


This assignment does not contain any printable content.

4.09 Misrepresenting Data with Graphical Displays


Misrepresenting Data with Graphical Displays
It is crucial to be aware that there are many ways that graphical displays can be manipulated and
edited to misrepresent data. Therefore, it is important to look at all of the key areas of graphical
displays. Specifically, the parts we can look at are axes, labels, data (i.e., is any data missing?),
perspective (i.e., is the graph 2D or 3D?), and overall construction. These are common aspects that
can be manipulated to change the way a graph looks, and make the data appear to illustrate a
conclusion that in reality is false. Take note: you can never be too careful when it comes to
analyzing graphical information!

Below are just a few of the ways that graphical displays can misrepresented.

Changing the Scale of Axis Labels

One of the most common examples is the following:

This graph shows the number of admissions per year for three universities. Do you notice
anything wrong with this graph? Compare it to the following:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


The second graph displays the same data. What's the difference between the two graphs?

In the first graph, the differences between the universities' admissions appear to be greater
than they do in the second graph. The reason is that the vertical scale does not start at
zero. This is known as truncating, which exaggerates the differences between the three
universities.

Omitting Axis Labels or Units

Another example of graphs that can misrepresent data are graphs that omit labels or units,
such as the pie chart displayed below.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Without the percentages for each of the sections of the pie chart, it is difficult to determine
which degree program is the largest, and which is the smallest.

Using a Two-Dimensional Figure to Represent a One-Dimensional Measurement

Another common source of misrepresentation is the use of a two-dimensional figure to


represent a one-dimensional measurement. For example, suppose we have data showing
the average healthcare spending per person in the U.S. increased from $146
in 1960
to $9, 532
in 2014
. If we use a bar graph, it might look like this:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


If instead of using a bar graph, we use a two-dimensional figure such as a circle or a
rectangle, or a three-dimensional figure such as a cylinder or a rectangular prism, we can
create the false impression of greater-than-real differences. The reason is that our eyes
see area (in the case of two-dimensional figures) and volume (in the case of three-
dimensional figures). This distorts the true differences we are trying to illustrate.

4.10 Vocabulary Game

4.11 Flashcards
Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.
4.11 Flashcards
Module 4 Flashcards

Term Definition

4.12 Review Game: Descriptive Statistics for a Single Variable

4.16 Module 4 Review Test

4.17 Module Feedback

You might also like