0% found this document useful (0 votes)
6 views38 pages

Module 5 Descriptive Statistics for Two Variables

Module 5 focuses on descriptive statistics for analyzing two variables, covering learning objectives such as classifying data analysis situations, identifying graphical displays, and summarizing distributions. It explains the roles of explanatory and response variables, along with appropriate graphical displays like two-way tables, side-by-side box plots, and scatterplots for different types of data. Additionally, it emphasizes the importance of conditional percentages in two-way frequency tables for understanding relationships between categorical variables.

Uploaded by

phillylpm
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views38 pages

Module 5 Descriptive Statistics for Two Variables

Module 5 focuses on descriptive statistics for analyzing two variables, covering learning objectives such as classifying data analysis situations, identifying graphical displays, and summarizing distributions. It explains the roles of explanatory and response variables, along with appropriate graphical displays like two-way tables, side-by-side box plots, and scatterplots for different types of data. Additionally, it emphasizes the importance of conditional percentages in two-way frequency tables for understanding relationships between categorical variables.

Uploaded by

phillylpm
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Module 5: Descriptive Statistics for Two Variables

Module 5: Descriptive Statistics for Two Variables

5.01 Learning Objectives


Module 5: Learning Objectives
After completing this module, you should be able to:

1. Classify a data analysis situation according to the role-type classification


2. Identify the appropriate graphical display for a given classification in different contexts
3. Identify the appropriate numerical measures for a given classification in different contexts
4. Compare conditional percentages in a two-way table
5. Summarize distributions of two variables
6. Determine the relationship between two quantitative variables, when given a graph
7. Describe the overall pattern and the striking deviations in a graph of two variables

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


5.02 Explanatory and Response Variables
Explanatory and Response Variables

Career Connections

Explanatory and Response Variables

Whenever you need to dig deeper to analyze data, you will likely need to
measure units of something that might be the cause and units of something
that might be the effect.

For example, suppose you are managing a security company that patrols
warehouses at night, and you suspect that the higher absenteeism in winter
is due to the lower temperatures. You can test your theory by gathering two
types of data. You can chart how cold a given night is and how many people came to work. If your
theory is correct, the data will show that as temperatures get colder, more employees will call in
sick.

Or in a manufacturing situation, you might suspect that a recent spate of defective merchandise
might be due to faulty goods from a particular supplier. You might track the number of items
returned that had that supplier's component to the number of items returned that had the same
component from another supplier. With enough data, comparing these numbers will shed light on
whether the supplier is the cause.

When one variable causes change in another, we call the first variable the explanatory variable (or
independent variable). The affected variable is called the response variable (or dependent variable).
In a randomized experiment, the researcher manipulates values of the explanatory variable and
measures the resulting changes in the response variable. The different values of the explanatory
variable are called treatments. An experimental unit is a single object or individual to be
measured.

Example
An insurance company wants to investigate whether taking a National Safety Council course in
workplace safety reduces the risk of on the job injuries. Four-hundred employees in both the

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


manufacturing and operations departments were recruited as participants. The employees were
divided randomly into two groups: one group will take the National Safety Council course, and the
other group will not receive additional safety training beyond the standard company training. Each
employee is then tracked over the course of the next two years and any employee that suffers an
on the job injury is recorded. At the end of the study, the insurance company counts the number of
employees in each group who suffered an on the job injury.
Identify the following values for this study: population, sample, explanatory variable, response
variable.
The population = Employees in the manufacturing and operations departments
The sample = The 400 employees who participated in the study
The explanatory variable = Whether the employee had additional safety training.
The response variable = Whether the employee incurred an injury on the job

Exercise

Identifying the Explanatory and Response Variables


1. A researcher wants to study the possible relationship between birth order and personality.
Identify the explanatory variable and the response variable. Enter in "explanatory" or "response"
for your answer.

a. Birth Order:
b. Effect on personality:
2. A recent study was conducted to determine if driving performance was influenced by texting.
Specifically, the research investigated how many seconds it takes for a driver to respond when a
leading car hits the brakes. Identify the explanatory variable and the response variable. Enter in
"explanatory" or "response" for your answer.

a. Response time measured in seconds:


b. Presence of distraction from texting:
3. The Smell & Taste Treatment and Research Foundation conducted a study to investigate
whether smell can affect learning. Subjects completed mazes multiple times while wearing masks.
They completed the pencil-and-paper mazes three times wearing floral-scented masks, and three
times with unscented masks. Participants were assigned at random to wear the floral mask during
the first three trials or during the last three trials. For each trial, researchers recorded the time it
took to complete the maze and the subject's impression of the mask's scent: positive, negative, or
neutral. Identify the explanatory variable and the response variable. Enter in "explanatory" or
"response" for your answer.

a. Scent:
b. Time it takes to complete the maze:
4. A study is being conducted comparing age and running speed for a sample of people of a
variety of ages. Which of the following best describes the relationship between age and running
speed?

a. Running speed is the explanatory variable, while age is the response variable.
b. Both age and running speed are explanatory variables.
c. Age is the explanatory variable, while running speed is the response variable.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


d. Both age and running speed are response variables.

5. A recent study examined the possible relationship between the use of a stand-up workstation
desk and productivity level. Identify the explanatory variable and the response variable. Enter in
"explanatory" or "response" for your answer.
a. Productivity:
b. Use of a stand-up workstation desk:

5.03 Graphical Displays for Two-Way Data


Graphical Displays for Two-way Data
When we analyze data that has two variables, if both variables are categorical C ( → C), we
display them in a two-way frequency table (or contingency table). If the explanatory variable is
categorical and the response variable is quantitative (C → Q ), we can use side-by-side box plots
to display them. If both variables are quantitative (Q → Q ), a scatterplot is usually the best choice
to display their relationship.

C → C: Two-way Tables
A two-way table, also known as a two-way frequency table or contingency table, is used to
show the relationship between two categorical variables (
C → C); the rows show the categories of one variable, and the columns show the
categories of the other variable.

In the table above, the cells in yellow show joint frequencies. These represent the total
number of instances that fall in both the corresponding row and header. For example, data
in the "Male" row and "With Autism" column counts the number of males with autism.

The data in the green cells show marginal frequencies. These are equal to the sum of the
number of individuals in the corresponding row or column. For example, data in the "Totals"
column and "Female" row shows the total number of females in the study. It may be helpful
to remember that marginal frequencies appear in the margins of the table.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


The bottom, right cell (in both the "Totals" column and the "Totals" row) measures the total
number of individuals in the study.

C → Q: Side-by-side Box Plots


Side-by-side Box Plots Review

When two-variables—one categorical, explanatory variable and one quantitative, response


variable—are to be displayed, side-by-side box plots are an effective choice. Both box plots
are displayed on the same graph. The side-by-side box plot shows a graphical comparison
for the two groups.

Vertical Side-by-side Box Plot

Example:

In this example, gender is a categorical variable and height is a quantitative variable.


Gender is also the explanatory variable and height is the response variable. These side-by-
side box plots show that, in general, males are about six inches taller than females. There
is about the same amount of variation within the heights of each gender.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Horizontal Side-by-Side Box Plot

Side-by-side box plots do not need to be presented vertically. Here is an example of a


horizontal side by side boxplot that compares data from two different stocks:

Here, the categorical, explanatory variable (Stock A and Stock B) are listed above and
below one another. The quantitative, response variable is displayed on a horizontal axis.

Q → Q: Scatterplots
The relationship between two variables that are both quantitative can be displayed in a
scatterplot.

A scatterplot is a graphical display on the coordinate plane that typically shows the
explanatory variable on the x-axis and the response variable on the y-axis. The values of
x and y for each subject are represented by a point on the scatterplot. As we've seen
earlier, every point on a coordinate plane can be represented by an ordered pair, (x, y).
Here, the x-value is typically the explanatory variable's value for a piece of data, and they-
value is the corresponding value for the response variable. A simple way to remember this
fact is that the term "explanatory" has an "x" in it.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


The points in this scatterplot show a trend: as waist circumference increases, arm
circumference increases. The relationship between the two variables is referred to as the
correlation.

Scatterplots will be discussed in more detail later in this module.

Exercise: Graphical Displays for Two Variables


Two-way Tables: Suppose we want to analyze the number of connection attempts to a server
that are suspicious connections. The results are summarized in the table below.

Outside Connection Type


Safe Connection Suspicious Connection
Server Web Server 60 23
Type Email Server 67 6
Side-by-side Box Plots:
Scatterplots:

Mini Case Study:


10. A network security firm is hired by a small finance company to perform comprehensive
information security services. As part of the service, the security manager analyzes the network
security settings versus the login errors received from illegitimate users. The security manager is
interested in deciphering the relationship between the network segment's security level (safe or
unsafe) and the number of illegitimate user login errors. For this analysis: (for the following

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


questions enter the letter that corresponds with your answer)

a. Which of the following is the explanatory variable?

a. Network segment
b. Network security level
c. Illegitimate user login errors

b. The explanatory variable is a quantitative variable. True or False?

a. True
b. False

c. Which of the following is the response variable?

a. Network segment
b. Network security level
c. Illegitimate user login errors

d. The response variable is categorical. True or False?

a. True
b. False

e. What graphical display should be used to show the results of the study?

a. Side by side boxplots


b. Bar chart
c. Scatterplot
d. None of the above

Exercise #5
11. When analyzing a possible relationship for two-variable data, if both variables are categorical,
what is the most appropriate choice to display the data?

a. Side-by-side boxplots
b. Scatterplot
c. Bar chart
d. Two-way frequency table
e. Histogram

12. When working with two-variable data, if the explanatory variable is categorical and the
response variable is quantitative, what is the most appropriate choice to display the data?

a. Side-by-side boxplots
b. Scatterplot

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


c. Bar chart
d. Two-way frequency table
e. Histogram
13. When working with two-variable data, if both variables are quantitative, what is the most
appropriate choice to display the data?

a. Side-by-side boxplots
b. Scatterplot
c. Bar chart
d. Two-way frequency table
e. Histogram

For the following two questions, refer to the boxplots below. The statistics refer to the difference in
salary for two groups of employees in different regions.

14. Which group has the greater percentage of employees with a difference in salary of less than
0?
a. Group A
b. Group B
c. Cannot be determined

15. Which group has more employees?

a. Group A
b. Group B
c. Cannot be determined

5.04 Numerical Analysis for C → C


Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.
5.04 Numerical Analysis for C → C
Numerical Analysis for C →C
Career Connections

Student STEM Career Interest

In education, people are worried about not being able to adequately


address future employment needs in science, technology, engineering, and
mathematics (called STEM) disciplines (see Sass, 2015 for a discussion on
this topic). This problem is often referred to as a shortage in the STEM
pipeline. So we in education need to do more to help our students engage
in STEM disciplines. There's a handy statistical tool that can help you track
this in your classroom: a two-way table.

As you'll see in this lesson, a two-way table is a way to compare two categorical variables. For
example, you could compare interest in a STEM career (interested or not interested) to the gender
of the student (female or male) using a two-way table. This could help you keep track of student
interest by gender. Why is it important to keep track of this across genders? Because females tend
to choose non-STEM careers even when they score as well or better than their male counterparts
on standardized tests. By encouraging and helping your female students find a STEM career, you
can help address the needs of the future! Two-way tables can help you do this anytime you're
comparing two categorical variables.

This course examines three possible situations with two variable data:

1. Both variables are categorical (C → C)


2. One categorical, explanatory variable and one quantitative, response variable (C → Q)
3. Both variables are quantitative (Q → Q).

We will examine each of these three situations and discuss methods of numerical analysis for each.
On this page, we will be examining the situation in which both variables are categorical (C → C).

Conditional Percentages in a Two-way Frequency Table

If both variables are categorical, a two-way frequency table is used to display the data. The
following example of a two-way table shows the association between perceived stress level
and heart attacks for a group of executives. In this example, perceived stress level is the
explanatory variable and heart attack is the response variable.

Heart Attacks
(Response
Variable)
Has Has Not
Had a Had a
Total
Heart Heart
Attack Attack

214 2462 2676


Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.
Perceived Stress Level Low
High
214
667 2462
2671 2676
3338
(Explanatory Variable)
Total 881 5133 6014

To analyze this contingency table, we can calculate proportions for each cell in the table. If
we want to examine how the outcome of the response variable is explained by the value of
the explanatory variable, we calculate the relative frequencies, or conditional percentages,
for each row of the table. The reason we calculate the percentages for each row of the
table is because the explanatory variable's values are found in the first two rows of the
table. If the table had been constructed with the explanatory variable's values in the first
two columns of the table, we would need to calculate the percentages for each column of
the table instead.

If we want to find the conditional percentages, simply divide the number of instances by the
total number of individuals in the corresponding explanatory category.

Example: Association between perceived stress level and taking an early retirement

For example, examine the table below that illustrates the association between perceived
stress level and the rate of taking an early retirement for a population of executives.

Rate of Retiring Early


Did Not Take
Took Early Retirement Early Total
Retirement
214 2462 2676
Low 2676
= 2676
= 2676
=
Perceived
Stress
8% 90% 100%
667 2671 3338
3338 = 3338 = 3338 =
Level
High
20% 80% 100%
881 5133 6014
Total 6014
= 6014
= 6014
=
14. 6% 85. 4% 100%
If the explanatory variables are each in their own row, be sure the proportions in each row
sum to 100%. If the explanatory variables are each in their own column, be sure the
proportions in each column sum to 100%.

Here, the proportions in each row sum to100%. From this analysis, we can see that the
group of executives who reported feeling a low level of perceived stress had an early
retirement rate of 8%, while the group of executives who reported feeling a high level of
perceived stress had an early retirement rate of 20%, more than twice the rate of those
who reported a low level of perceived stress.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Conditional Percentages and Relative Frequency in a Two-way Table:

The frequency counts in each cell of the table are the joint frequencies. The totals in each row and
column are the marginal frequencies.

There are several ways we can analyze the data presented in this table. If we calculate the
percentage that each cell is of the total, the results are called relative frequencies. When the
relative frequencies are calculated from the row total or the column total, they are called conditional
percentages.

The table below shows frequency counts for preferred exercise activities for 50 adults—25 women
and 25 men.

Exercise Program
Yoga Class Weightlifting Dance Exercise Total
Women 11 3 11 25
Men 4 16 5 25
Total 15 19 16 50
We typically find the conditional percentages based on the explanatory variable. However, since the
explanatory variable could be in either the row or the columns, it's important to be mindful of the
orientation of the explanatory variable we are calculating.

Examine the tables below to see these different situations.

Relative Frequencies by Row

Each row is a different gender. If we are trying to see if gender influences the choice of exercise
program, then gender is the explanatory variable. In this case, we are calculating the relative
frequency by rows; that is, we are calculating the relative frequency by gender. To determine
relative frequency for women, we divide the data in the top row by the total number of women. To
determine the relative frequency for men, we divide the data in the second row by the total number
of men. The percentages obtained are called conditional row percentages.

Exercise Program
Yoga Class Weightlifting Dance Exercise Total
Women 0. 44 0. 12 0. 44 1. 00
Men 0. 16 0. 64 0. 20 1. 00
Total 0. 30 0. 38 0. 32 1. 00
Remember that decimal values can be converted to a percentage by moving the decimal to the
right two places - so . 44 means that you are looking at 44%. The table above shows that among
women, yoga class (44%) and dance exercise (44%) are the preferred exercise programs, while
among men, weightlifting (64%) is the preferred choice. As the explanatory variables each have
their own row, be sure that each row sums to 100%.

Relative Frequencies by Column

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Each column is a different exercise program. If we are trying to determine how each exercise
program is appealing to different genders, then the exercise program becomes the explanatory
variable. In this case, the explanatory variable is in the columns, so we will be calculating the
relative frequency by columns, that is, we are calculating the relative frequency by exercise
program. To determine relative frequency for each cell, we divide the data by the corresponding
column's total number of individuals. The percentages obtained are called conditional column
percentages.

Exercise Program
Yoga Class Weightlifting Dance Exercise Total
Women 0. 73 0. 16 0. 69 0. 50
Men 0. 27 0. 84 0. 31 0. 50
Total 1. 00 1. 00 1. 00 1. 00
From this analysis, we can conclude that, based on this sample, we can expect that a yoga class
3 1 (27%) men. We can also conclude that about
would have approximately (73%) women and
4 4
4 (84%) of weightlifters are men, and that 2 (69%) of dance exercisers are women.
5 3

Overall Percentages

We can also calculate percentages for the whole table. Here, each cell's data is divided by the
total number of individuals. By calculating the overall percentages for the whole table, we are
determining the relative frequency of each combination of activities: for example, what percentage
of people are women who participated in a yoga class?

Exercise Program
Yoga Class Weightlifting Dance Exercise Total
Women 0. 22 0. 06 0. 22 0. 50
Men 0. 08 0. 32 0. 10 0. 50
Total 0. 30 0. 38 0. 32 1. 00
Analysis of this table shows that while the three exercise programs have roughly the same overall
level of interest, there is a big difference between men and women in their preferred exercise
programs. The bottom right cell should read 1. 00. It reads 1. 00 because the rest of the bottom
row, as well as the rest of the right column, sums to 1. 00.

Career Connections

Conditional Percentages in a Two-Way Table

Conditional percentages in a two-way table are used when measuring


categorical data. Although we think of business decisions as primarily about
numbers, many business decisions require sorting through categorical data.

For example, a company looking to buy new hardware might wish to measure the effectiveness of

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


its old desktop computers. To that end, management might create a two-
way table that measures why desktop computers were disposed of, such as
"suffered from a malware attack," "hard drive malfunction" and
"outdated/slow." The number (or relative frequency) of machines disposed
of for each reason is recorded in the columns. The table would plot the
number (or relative frequency) of machines from each manufacturer in each
of the rows. The table would give a good sense of which issues affected
each manufacturer so the company can make the best decisions about
what computers to buy now.

5.05 Numerical Analysis for C → Q and Q → Q


Numerical Analysis for C → Q and Q → Q
On the previous page, we examined the situation in which both variables are categorical C ( → C).
On this page, we will be examining the situations in which there is one categorical (explanatory)
variable and one quantitative (response) variable (
C → Q) as well as the situation in which both variables are quantitative (
Q → Q).

C → Q: Five-number Summary
Five-number Summary Review

If the explanatory variable is categorical and the response variable is quantitative, we can
use descriptive statistics, namely the five-number summary, for the quantitative variable,
and compare the statistics for each of the categories. You might have already guessed this
since you know that the side-by-side box plots are used for the C → Q classification.

Let's look at an example. Consider a study on two different stocks and their daily returns.
Total daily return (as a percentage, positive or negative) is calculated for two stocks: Stock
A and Stock B. Therefore, the daily return is the quantitative response variable, and the
stock (Stock A/Stock B) is the categorical explanatory variable.

Total Daily Return


Stock A −0. 27 1. 42 2. 12 2. 24 −0. 51 −0. 53 −1. 19 0. 33 1. 19 0. 12

Stock B −2. 29 −0. 52 −0. 11 −0. 35 −1. 47 −0. 45 −0. 83 −0. 13 0. 42 0. 78

We can calculate the five-number summary for each group.

Total Daily Return Minimum Q1 Median Q3 Maximum


Stock A −1. 19 −0. 51 0. 225 1. 42 2. 24
Stock B −2. 29 −0. 83 −0. 4 −0. 11 0. 78
Q → Q: Correlation Coefficient (r)

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Types of Relationships

If both variables are quantitative, a scatterplot can show whether or not there is a
relationship between the variables. There are two types of relationships that can exist
between two quantitative variables — a linear relationship or a curvilinear relationship,
which is sometimes called a nonlinear relationship. A linear relationship is a relationship
between two quantitative variables in which the points in the scatterplot follow a pattern that
forms a reasonably straight line, as in the following scatterplot of the total number of bugs
reported after the launch of a new app vs. the number of days since the app was launched.

A curvilinear relationship, on the other hand, is a relationship between two quantitative


variables in which the points in the scatterplot follow a pattern that is more curved than
straight. The following scatterplot shows a curvilinear relationship between the variables.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Extracting Information From a Scatterplot

We can make some observations about the relationship between the variables number of
days since the app was launched and total number of bugs reported after the launch of a
new app by inspecting the scatterplot of total number of bugs reported after the launch of a
new app vs. number of days since the app was launched above.

First, the explanatory variable is number of days since the app was launched and the
response variable is total number of bugs reported after the launch of a new app, since
number of days since the app was launched is on the x-axis and total number of bugs
reported after the launch of a new app is on the y-axis.

We can also see that the variables move in the same direction; in other words, as the
number of days since the app was launched increases, the total number of bugs reported
after the launch of a new app also increases and as the number of days since the app was
launched decreases, total number of bugs reported after the launch of a new app also
decreases.

When there is a linear relationship between two quantitative variables and both variables
change in the same direction, we say there is a positive correlation between the variables.
On a scatterplot, you can recognize a positive correlation by noticing that as you move from
left to right on the graph, the points move upward.

A positive correlation between two quantitative variables is in contrast to a negative


correlation, which is characterized by the two variables moving in opposite directions; that
is, as one variable increases, the other variable decreases. On a scatterplot, you can
recognize a negative correlation by noticing that as you move from left to right on the
graph, the points move downward.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


If you have ever looked at people's productivity at work, you've seen a negative correlation
in action. For example, among a group of sales people, as the amount of time they miss
from work increases, the amount of their sales decreases. Conversely, the less time they
miss from work, the more they sell.

Correlation Coefficient

Notice that the scatterplot of total number of bugs reported after the launch of a new app
vs. number of days since the app was launched demonstrates a linear relationship. Even
though there are a few data points that do not fit the linear trend, the majority of the data
points fit the trend "well enough" for the relationship to be linear.

This statement leads to the question: How do we know if the points on a scatterplot of two
variables fit the pattern "well enough" to conclude there is a linear relationship between the
two variables? To answer this question, we use a quantitative measure that tells us how
closely the data values in a scatterplot fall to a straight line; a measure called the
correlation coefficient. The correlation coefficient is represented by r and is a measure of
the strength and direction of the linear relationship.

The correlation coefficient (r) is a number that falls somewhere from −1 to 1. The closer to
0 that the correlation coefficient (r) is, the weaker the linear relationship.

A correlation coefficient (r) at or near −1 represents a strong negative linear


correlation.

A correlation coefficient (r) at or near 1 represents a strong positive linear correlation.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


We will explore scatterplots and correlation coefficients in greater detail on a later page in
this course.

Note — the correlation coefficient (r) gives information on the strength and direction of the
relationship only if the relationship is linear!

Exercise

One Variable Categorical, One Variable Quantitative


For the following questions, please enter the letter in the answer box that corresponds with your
answer choice.

The side-by-side box plot shows a graphical comparison for the two groups. Using this comparison,
answer the following questions.

1. Which stock has the greater median daily return?

A. Stock A
B. Stock B

2. Which group has a larger Interquartile range?

A. Stock A
B. Stock B

3. Which group has less variation in its data?

A. Stock A
B. Stock B

4. What is the five-number summary for Stock A?

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


a. Minimum:
b. Q1:
c. Median:
d. Q3:
e. Maximum:

5. Based on the data in the following table, calculate the Five-number Summary.
Tasks Completed
Population Sample 178 100 153 168 111 91 105 110
a. Minimum:
b. Q1:
c. Median:
d. Q3:
e. Maximum:

Both Variables Quantitative


6. A scatterplot is useful for which type of data?

a. Both variables are categorical.


b. One variable is categorical, one variable is quantitative.
c. Both variables are quantitative.
d. One variable is discrete, one variable is continuous.

7. True or False? A scatterplot always shows the explanatory variable on the horizontal, or x-axis.

a. True
b. False

8. The strength of the correlation between two variables can be measured by which statistic?

a. Relationship factor
b. Correlation coefficient (r)
c. Strength test
d. Trending statistic

Review Checkpoint
To test your understanding of the content presented in this assignment, please click on the
Question icon below. Click your selected response to see feedback displayed below it. If you have
trouble answering, you are always free to return to this or any assignment to re-read the material.

1. A company is studying the effectiveness of two different production processes (Process A and
Process B) for a week. The results are measured by counting the number of products without
defects and with defects. Which numerical measure could be used to analyze the data?

a. Correlation coefficient (r)

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Incorrect. Try again.

b. Five-number summary

Incorrect. Try again.

c. Median

Incorrect. Try again.

d. Conditional percentages

Correct. The answer is d. Both variables are categorical (C → C) so we will use a two-way table.
Therefore, our numerical measure will be conditional percentages.

2. The company also wants to measure the time each process takes. The company measured the
time it took to make each product as it came off of the assembly line. Which numerical measure
would present the maximum and minimum amount of time that it took for products to be made using
each process system?

a. Relative Frequency

Incorrect. Try again.

b. Five-number Summary

Correct. The answer is b. In this study one variable is categorical, and the other is quantitative
(C → Q). The five-number summary will show five important statistical values: minimum,
maximum, first quartile, median, and third quartile. Therefore, the five-number summary would be
the best choice to look at the minimum and maximum values for both processes.

c. Correlation

Incorrect. Try again.

d. Conditional percentages

Incorrect. Try again.

3. A marketer was interested in how age affects responses to images of healthy food. He sampled
a population ranging in age from 10 years old to 70 years old. Subjects were asked to complete a
simple survey, rating how much they would like to eat the food shown in 10 images. The rating went
from 1 (not at all) to 10 (very much). The food depicted was all food considered healthy. Which
numerical measure could be used to analyze the strength of the relationship between age and
response to images of healthy food?

a. Correlation coefficient (r)

Correct. The answer is a. In this study both variables are quantitative(Q → Q) that form paired
data. The strength of a correlation can be measured by calculating the correlation coefficient (r).

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


b. Conditional percentages

Incorrect. Try again.

c. Mean

Incorrect. Try again.

d. Five-number Summary

Incorrect. Try again.

4. A car company is comparing the miles its trucks can go before needing repair to that of one of its
rivals. The researchers want to compare the middle 50% of the data for both groups. Which
numerical measure would identify the two values that 50% of the data falls between for both
groups?

a. Median

Incorrect. Try again.

b. Five-number summary

Correct. The answer is b. In this study one variable is categorical and the other is quantitative
(C → Q). The five-number summary will show five important statistical values: minimum,
maximum, first quartile, median, and third quartile. 50% of the data falls between the first(Q1) and
third(Q3) quartile, so the researchers should look at the five-number summary to determine the
values for Q1 and Q3 for both groups, as these values define the middle50% of the data.

c. Mode

Incorrect. Try again.

d. Joint Frequencies

Incorrect. Try again.

5. The safety and quality assurance department recently implemented a new product testing
program. The department was interested in determining if this new program decreased a products's
risk of stress related failure within two years. Group A was exposed to the new product testing
program, while the Group B received the traditional testing program. Which numerical measure
could be used to analyze the data?

Stress Related Failure within 2 years


Yes No Total
Group A 11 39 50
Group B 29 21 50
Total 40 60 100
a. Marginal Frequency

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Incorrect. Try again.

b. Relative Frequency

Incorrect. Try again.

c. Joint Frequency

Incorrect. Try again.

d. Conditional percentages

Correct. The answer is d. In this study both variables are categorical(C → C), so a two-way table
is used to present the data. Therefore, our numerical measure will be conditional percentages.

5.06 Relationships Between Two Variables on a Scatterplot (Q


→ Q)
Relationships Between Two Variables on a Scatterplot (Q → Q)
Career Connections

Looking for a Relationship between Age and Adoption of Technology

You may have heard that the speed at which someone adopts new
technology is dependent on his or her age. Said another way, younger
people more readily engage in new technology while older people are less
likely to engage in new technology. There is some truth to this, but the issue
is a lot more complicated. How could you look at data on these two
quantitative variables (age and likelihood of adopting new technology) to
see if there is a relationship? By using a tool called a scatterplot!

Scatterplots are one of the easiest tools we have for comparing two quantitative variables. As you'll
see in this lesson, scatterplots give us a visual depiction of the relationship (or lack of relationship)
between two variables. We use a scatterplot for two quantitative variables because we graph the
relationship on an x-y-plane, similar to what you saw in Module 3. Since we do this on thex-y-
plane, the two variables we're looking at have to be something we can measure, which is exactly
why we use a scatterplot for two quantitative variables.

Other examples of relationships in IT you might want to look at using a scatterplot include:

Speed of computer processors over the years


Number of help requests over the days following a new update
Total budget of an IT department compared to the total number of company employees

Remember, if you are looking at two quantitative variables, a scatterplot is a great tool to see how
the two variables compare to one another.

Summarizing Distributions of Two Variables

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


When we have two quantitative variables that are from paired data (that is, eachx-value is paired
with a particular y-value), we can use a scatterplot to display the data. The values of the
explanatory variable appear on the x-axis, and the values of the response variable appear on they-
axis.

Each pair of (x, y) values appears as a point on the scatterplot, and when we see a completed
scatterplot, it forms a picture of the data that shows us the relationship between the variables. When
we look at the picture, we look for an overall pattern and whether there are any deviations (outliers)
from the pattern. We also can describe the scatterplot by the direction, the form, and the strength of
the relationship.

There are four examples of scatterplots, representing (


Q → Q): positive correlation, negative correlation, no correlation, non-linear relationship

Positive Correlation

Scatterplot (a) shows a positive correlation between the variables because as thex-
variable increases, the y-variable increases.

Negative Correlation

Scatterplot (b) shows a negative correlation because as thex-variable increases, the y-


variable decreases.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


No Correlation

Scatterplot (c) shows no correlation; there is no apparent overall trend between the two
variables.

Non-Linear Relationship

Scatterplot (d) shows a nonlinear, or curvilinear, relationship.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Scatterplots can also be characterized by how closely the points follow a clear form. For example, if
all of the points line up in a perfectly straight line in a positive direction (up and to the right), it is said
to be a perfectly linear positive correlation. If the points are close to a straight line in the positive
direction, but do not form a perfectly straight line, we say it is a strong positive correlation. If all of
the points line up in a perfectly straight line in a negative direction (down and to the right), it is said
to be a perfectly linear negative correlation. If the points are close to a straight line in the negative
direction, but do not form a perfectly straight line, we say it is a strong negative correlation.

Describing the Relationship Between Two Quantitative Variables

The relationship between two quantitative variables (


Q → Q) can be described by looking at a scatterplot of the two variables. We use three
characteristics to describe the relationship: direction, form, and strength.

Direction

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


If a scatterplot shows a pattern of points that increase from the lower left corner of the
graph to the upper right corner, we say that there is a positive correlation between the two
variables: when the x-variable increases, the y-variable increases. When the pattern goes
from the upper left corner of the graph to the lower right corner, we say that there is a
negative correlation between the two variables: when the x-variable increases, the y-
variable decreases.

Form

If a scatterplot has a pattern of points that form a reasonably straight line, we describe it as
linear. If the points form a pattern that is more curved than straight, we say it is nonlinear,
or curvilinear.

Strength

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


The pattern of points tells us about the strength of the correlation between the variables. If
the points form a tightly grouped pattern, we say there is a strong correlation between the
variables. If the points are loosely scattered and are not tightly grouped, we say there is a
weak correlation or no correlation between the variables.

Putting it All Together

Here are some examples of scatterplots with descriptions using all three characteristics to
describe the relationship between the x-variable and the y-variable.

Examples

1. Strong positive linear correlation

2. Strong negative linear correlation

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


3. Very weak/no correlation

4. Nonlinear (curvilinear) association

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


Exercise

Summarizing Distributions of Two Variables


Match each scatterplot below with the most appropriate description of its correlation.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


1. A perfectly linear negative correlation.
2. No association or relationship between the variables.
3. A strong, but not perfectly linear, positive correlation.
4. A strong, but not perfectly linear, negative correlation.
5. A perfectly linear positive correlation.
6. A non-linear relationship.
Describing the Relationship Between Two Quantitative Variables

5.07 The Scatterplot


The Scatterplot
As previously discussed, a scatterplot is a type of graph on a coordinate plane. A scatterplot helps
to show any potential relationships between two quantitative variables, (Q → Q ), represented by
their x- and y-coordinates on a coordinate plane. Data points are plotted as dots on the coordinate
plane, and the concentration, dispersion, or overall trend of these dots shows if there is relationship
between the variables, and what that relationship looks like.

To construct a scatterplot to see any potential relationship between variables, follow the instructions

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


below:

Step 1. Collect data

The data being collected should have two variables. These variables will correspond to thex-axis
and y-axis of your scatterplot. Determine the explanatory variable (a suspected "cause") and the
response variable (a suspected "effect").

For example, a company that specializes in hand-crafted furniture conducts a study to determine a
possible relationship between the size of the bookcases it makes and the average time it takes its
carpenters to assemble the bookcase—that is, the company is investigating if the total board length
(measured in feet) contained in a single bookcase affects the average time (measured in minutes)
needed to assemble that bookcase. Therefore, board length is the explanatory variable on the x-
axis and average assembly time is the response variable on the y-axis. Here is the data that was
collected:

Styles of Board Length in Bookcase Average Time of


Bookcases (ft) Assembly
Product #1 28 210
Product #2 20 185
Product #3 25 207
Product #4 29 226
Product #5 28 236
Product #6 31 235
Product#7 31 220
Product #8 33 241
Product #9 21 179
Product #10 21 165

Step 2. Plot the data points on the coordinate plane

Each dot on the scatterplot represents one style of bookcase, a particular product. The product's
board length and average assembly time can be seen as ordered pairs. The product's board length
measurement is the x-value and the assembly time is the y-value. Here are these points plotted on
a coordinate plane:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


For a scatterplot, the origin should be in view of the displayed graph whenever possible. However, if
including the origin in the scatterplot causes the displayed graph to be too large to fit on the page,
cutting the graph down to just the relevant area (where the scatterplot points are) is a common
practice.

Here, that means we can display 15 through 35 on the x-axis (all of these bookcases' board
lengths fall in this range), and we can display 100 through 260 on the y-axis (all of the bookcases'
average assembly times fall in this range). It is also best when creating a scatterplot that the x-axis
and y-axis are both labeled with appropriate units. A scatterplot should have a title over the top of
the graph. Once these changes are made, the scatterplot appears as follows:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


5.08 The Effects of Outliers on the Analysis of Two
Quantitative Variables
The Effects of Outliers on the Analysis of Two Quantitative Variables
When analyzing scatterplot displays of two quantitative variables, (Q → Q), we must be aware of
the possibility that one or more outliers can dramatically influence the results. An outlier in a
scatterplot is a point that does not fit the overall trend of the other points. It is usually far away from
the other points and is easy to spot.

There may be more than one outlier, but if several points are grouped together away from the
majority of points, we call them a cluster, not outliers. Consider the following scatterplot:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


There is one outlier in this data set: the data point in the upper left corner of the scatterplot. If it is
included in the analysis of the data, it will skew the measures of correlation. If it is not included,
there is an almost perfect strong positive linear correlation between the explanatory variable (shoe
size) and the response variable (height). This example illustrates the importance of plotting the data
and looking for patterns when analyzing two quantitative variables.

Outliers and Correlation

The presence of an outlier may affect the correlation coefficient r, depending on the placement of
the outlier and the sample size of the data. If an outlier falls far from the regression line, it can
weaken the correlation and move the correlation coefficient r closer to 0. If this outlier is removed,
there will be a stronger relationship between the two variables. On the other hand, if an outlier falls
near the regression line, it will have a diminished effect on the correlation, and may even strengthen
the correlation.

Notice that on the graphs below we have plotted the ordered pairs, as well as a line. This line is
called the "line of best fit" or the "regression line," and it will be discussed in more detail in a later
module. The line of best fit is the line that minimizes the distance between each point and the line
itself. While we can draw many lines on the scatterplot, the line of best fit is the line that is the
closest to all the points in the scatterplot. The closer the points are overall to the line of best fit, the
stronger the correlation (or linear relationship) will be. We are including it here in these graphs to
give you a better feel of the strength of the linear relationship.

Examine the graph below:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


This scatterplot has 21 data points, including the outlier that is above and to the left of the rest of
the data. There is a moderately strong positive correlation when this outlier is included. Once we
remove the outlier, however, the correlation becomes stronger.

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


The scatterplot above is the same as the earlier graph, but with the outlier removed. The20
remaining points have a very strong positive correlation. Note that by comparing the two graphs
(with the outlier and without the outlier), in the second graph, the points are overall closer to the line
of best fit, indicating a stronger correlation.

An outlier will have an even greater effect if the sample size is smaller. Compare the following two
graphs. First, with the outlier:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


The graph above has only six data points. Notice that there is a weak positive correlation. The
outlier (above and to the left of the rest of the data) has a dramatic effect on the correlation. Without
the outlier the correlation is much stronger:

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.


In the graph above, the five remaining data points have a very strong positive correlation. Removing
the outlier changed the correlation from weak to very strong. Now the points are much closer to the
line of best fit.

5.09 Flashcards
Module 5 Flashcards

Term Definition

5.10 Review Game

5.14 Module 5 Review Test

5.15 Module Feedback

Copyright © 2019 MindEdge Inc. All rights reserved. Duplication prohibited.

You might also like