Module 5 Descriptive Statistics for Two Variables
Module 5 Descriptive Statistics for Two Variables
Career Connections
Whenever you need to dig deeper to analyze data, you will likely need to
measure units of something that might be the cause and units of something
that might be the effect.
For example, suppose you are managing a security company that patrols
warehouses at night, and you suspect that the higher absenteeism in winter
is due to the lower temperatures. You can test your theory by gathering two
types of data. You can chart how cold a given night is and how many people came to work. If your
theory is correct, the data will show that as temperatures get colder, more employees will call in
sick.
Or in a manufacturing situation, you might suspect that a recent spate of defective merchandise
might be due to faulty goods from a particular supplier. You might track the number of items
returned that had that supplier's component to the number of items returned that had the same
component from another supplier. With enough data, comparing these numbers will shed light on
whether the supplier is the cause.
When one variable causes change in another, we call the first variable the explanatory variable (or
independent variable). The affected variable is called the response variable (or dependent variable).
In a randomized experiment, the researcher manipulates values of the explanatory variable and
measures the resulting changes in the response variable. The different values of the explanatory
variable are called treatments. An experimental unit is a single object or individual to be
measured.
Example
An insurance company wants to investigate whether taking a National Safety Council course in
workplace safety reduces the risk of on the job injuries. Four-hundred employees in both the
Exercise
a. Birth Order:
b. Effect on personality:
2. A recent study was conducted to determine if driving performance was influenced by texting.
Specifically, the research investigated how many seconds it takes for a driver to respond when a
leading car hits the brakes. Identify the explanatory variable and the response variable. Enter in
"explanatory" or "response" for your answer.
a. Scent:
b. Time it takes to complete the maze:
4. A study is being conducted comparing age and running speed for a sample of people of a
variety of ages. Which of the following best describes the relationship between age and running
speed?
a. Running speed is the explanatory variable, while age is the response variable.
b. Both age and running speed are explanatory variables.
c. Age is the explanatory variable, while running speed is the response variable.
5. A recent study examined the possible relationship between the use of a stand-up workstation
desk and productivity level. Identify the explanatory variable and the response variable. Enter in
"explanatory" or "response" for your answer.
a. Productivity:
b. Use of a stand-up workstation desk:
C → C: Two-way Tables
A two-way table, also known as a two-way frequency table or contingency table, is used to
show the relationship between two categorical variables (
C → C); the rows show the categories of one variable, and the columns show the
categories of the other variable.
In the table above, the cells in yellow show joint frequencies. These represent the total
number of instances that fall in both the corresponding row and header. For example, data
in the "Male" row and "With Autism" column counts the number of males with autism.
The data in the green cells show marginal frequencies. These are equal to the sum of the
number of individuals in the corresponding row or column. For example, data in the "Totals"
column and "Female" row shows the total number of females in the study. It may be helpful
to remember that marginal frequencies appear in the margins of the table.
Example:
Here, the categorical, explanatory variable (Stock A and Stock B) are listed above and
below one another. The quantitative, response variable is displayed on a horizontal axis.
Q → Q: Scatterplots
The relationship between two variables that are both quantitative can be displayed in a
scatterplot.
A scatterplot is a graphical display on the coordinate plane that typically shows the
explanatory variable on the x-axis and the response variable on the y-axis. The values of
x and y for each subject are represented by a point on the scatterplot. As we've seen
earlier, every point on a coordinate plane can be represented by an ordered pair, (x, y).
Here, the x-value is typically the explanatory variable's value for a piece of data, and they-
value is the corresponding value for the response variable. A simple way to remember this
fact is that the term "explanatory" has an "x" in it.
a. Network segment
b. Network security level
c. Illegitimate user login errors
a. True
b. False
a. Network segment
b. Network security level
c. Illegitimate user login errors
a. True
b. False
e. What graphical display should be used to show the results of the study?
Exercise #5
11. When analyzing a possible relationship for two-variable data, if both variables are categorical,
what is the most appropriate choice to display the data?
a. Side-by-side boxplots
b. Scatterplot
c. Bar chart
d. Two-way frequency table
e. Histogram
12. When working with two-variable data, if the explanatory variable is categorical and the
response variable is quantitative, what is the most appropriate choice to display the data?
a. Side-by-side boxplots
b. Scatterplot
a. Side-by-side boxplots
b. Scatterplot
c. Bar chart
d. Two-way frequency table
e. Histogram
For the following two questions, refer to the boxplots below. The statistics refer to the difference in
salary for two groups of employees in different regions.
14. Which group has the greater percentage of employees with a difference in salary of less than
0?
a. Group A
b. Group B
c. Cannot be determined
a. Group A
b. Group B
c. Cannot be determined
As you'll see in this lesson, a two-way table is a way to compare two categorical variables. For
example, you could compare interest in a STEM career (interested or not interested) to the gender
of the student (female or male) using a two-way table. This could help you keep track of student
interest by gender. Why is it important to keep track of this across genders? Because females tend
to choose non-STEM careers even when they score as well or better than their male counterparts
on standardized tests. By encouraging and helping your female students find a STEM career, you
can help address the needs of the future! Two-way tables can help you do this anytime you're
comparing two categorical variables.
This course examines three possible situations with two variable data:
We will examine each of these three situations and discuss methods of numerical analysis for each.
On this page, we will be examining the situation in which both variables are categorical (C → C).
If both variables are categorical, a two-way frequency table is used to display the data. The
following example of a two-way table shows the association between perceived stress level
and heart attacks for a group of executives. In this example, perceived stress level is the
explanatory variable and heart attack is the response variable.
Heart Attacks
(Response
Variable)
Has Has Not
Had a Had a
Total
Heart Heart
Attack Attack
To analyze this contingency table, we can calculate proportions for each cell in the table. If
we want to examine how the outcome of the response variable is explained by the value of
the explanatory variable, we calculate the relative frequencies, or conditional percentages,
for each row of the table. The reason we calculate the percentages for each row of the
table is because the explanatory variable's values are found in the first two rows of the
table. If the table had been constructed with the explanatory variable's values in the first
two columns of the table, we would need to calculate the percentages for each column of
the table instead.
If we want to find the conditional percentages, simply divide the number of instances by the
total number of individuals in the corresponding explanatory category.
Example: Association between perceived stress level and taking an early retirement
For example, examine the table below that illustrates the association between perceived
stress level and the rate of taking an early retirement for a population of executives.
Here, the proportions in each row sum to100%. From this analysis, we can see that the
group of executives who reported feeling a low level of perceived stress had an early
retirement rate of 8%, while the group of executives who reported feeling a high level of
perceived stress had an early retirement rate of 20%, more than twice the rate of those
who reported a low level of perceived stress.
The frequency counts in each cell of the table are the joint frequencies. The totals in each row and
column are the marginal frequencies.
There are several ways we can analyze the data presented in this table. If we calculate the
percentage that each cell is of the total, the results are called relative frequencies. When the
relative frequencies are calculated from the row total or the column total, they are called conditional
percentages.
The table below shows frequency counts for preferred exercise activities for 50 adults—25 women
and 25 men.
Exercise Program
Yoga Class Weightlifting Dance Exercise Total
Women 11 3 11 25
Men 4 16 5 25
Total 15 19 16 50
We typically find the conditional percentages based on the explanatory variable. However, since the
explanatory variable could be in either the row or the columns, it's important to be mindful of the
orientation of the explanatory variable we are calculating.
Each row is a different gender. If we are trying to see if gender influences the choice of exercise
program, then gender is the explanatory variable. In this case, we are calculating the relative
frequency by rows; that is, we are calculating the relative frequency by gender. To determine
relative frequency for women, we divide the data in the top row by the total number of women. To
determine the relative frequency for men, we divide the data in the second row by the total number
of men. The percentages obtained are called conditional row percentages.
Exercise Program
Yoga Class Weightlifting Dance Exercise Total
Women 0. 44 0. 12 0. 44 1. 00
Men 0. 16 0. 64 0. 20 1. 00
Total 0. 30 0. 38 0. 32 1. 00
Remember that decimal values can be converted to a percentage by moving the decimal to the
right two places - so . 44 means that you are looking at 44%. The table above shows that among
women, yoga class (44%) and dance exercise (44%) are the preferred exercise programs, while
among men, weightlifting (64%) is the preferred choice. As the explanatory variables each have
their own row, be sure that each row sums to 100%.
Exercise Program
Yoga Class Weightlifting Dance Exercise Total
Women 0. 73 0. 16 0. 69 0. 50
Men 0. 27 0. 84 0. 31 0. 50
Total 1. 00 1. 00 1. 00 1. 00
From this analysis, we can conclude that, based on this sample, we can expect that a yoga class
3 1 (27%) men. We can also conclude that about
would have approximately (73%) women and
4 4
4 (84%) of weightlifters are men, and that 2 (69%) of dance exercisers are women.
5 3
Overall Percentages
We can also calculate percentages for the whole table. Here, each cell's data is divided by the
total number of individuals. By calculating the overall percentages for the whole table, we are
determining the relative frequency of each combination of activities: for example, what percentage
of people are women who participated in a yoga class?
Exercise Program
Yoga Class Weightlifting Dance Exercise Total
Women 0. 22 0. 06 0. 22 0. 50
Men 0. 08 0. 32 0. 10 0. 50
Total 0. 30 0. 38 0. 32 1. 00
Analysis of this table shows that while the three exercise programs have roughly the same overall
level of interest, there is a big difference between men and women in their preferred exercise
programs. The bottom right cell should read 1. 00. It reads 1. 00 because the rest of the bottom
row, as well as the rest of the right column, sums to 1. 00.
Career Connections
For example, a company looking to buy new hardware might wish to measure the effectiveness of
C → Q: Five-number Summary
Five-number Summary Review
If the explanatory variable is categorical and the response variable is quantitative, we can
use descriptive statistics, namely the five-number summary, for the quantitative variable,
and compare the statistics for each of the categories. You might have already guessed this
since you know that the side-by-side box plots are used for the C → Q classification.
Let's look at an example. Consider a study on two different stocks and their daily returns.
Total daily return (as a percentage, positive or negative) is calculated for two stocks: Stock
A and Stock B. Therefore, the daily return is the quantitative response variable, and the
stock (Stock A/Stock B) is the categorical explanatory variable.
If both variables are quantitative, a scatterplot can show whether or not there is a
relationship between the variables. There are two types of relationships that can exist
between two quantitative variables — a linear relationship or a curvilinear relationship,
which is sometimes called a nonlinear relationship. A linear relationship is a relationship
between two quantitative variables in which the points in the scatterplot follow a pattern that
forms a reasonably straight line, as in the following scatterplot of the total number of bugs
reported after the launch of a new app vs. the number of days since the app was launched.
We can make some observations about the relationship between the variables number of
days since the app was launched and total number of bugs reported after the launch of a
new app by inspecting the scatterplot of total number of bugs reported after the launch of a
new app vs. number of days since the app was launched above.
First, the explanatory variable is number of days since the app was launched and the
response variable is total number of bugs reported after the launch of a new app, since
number of days since the app was launched is on the x-axis and total number of bugs
reported after the launch of a new app is on the y-axis.
We can also see that the variables move in the same direction; in other words, as the
number of days since the app was launched increases, the total number of bugs reported
after the launch of a new app also increases and as the number of days since the app was
launched decreases, total number of bugs reported after the launch of a new app also
decreases.
When there is a linear relationship between two quantitative variables and both variables
change in the same direction, we say there is a positive correlation between the variables.
On a scatterplot, you can recognize a positive correlation by noticing that as you move from
left to right on the graph, the points move upward.
Correlation Coefficient
Notice that the scatterplot of total number of bugs reported after the launch of a new app
vs. number of days since the app was launched demonstrates a linear relationship. Even
though there are a few data points that do not fit the linear trend, the majority of the data
points fit the trend "well enough" for the relationship to be linear.
This statement leads to the question: How do we know if the points on a scatterplot of two
variables fit the pattern "well enough" to conclude there is a linear relationship between the
two variables? To answer this question, we use a quantitative measure that tells us how
closely the data values in a scatterplot fall to a straight line; a measure called the
correlation coefficient. The correlation coefficient is represented by r and is a measure of
the strength and direction of the linear relationship.
The correlation coefficient (r) is a number that falls somewhere from −1 to 1. The closer to
0 that the correlation coefficient (r) is, the weaker the linear relationship.
Note — the correlation coefficient (r) gives information on the strength and direction of the
relationship only if the relationship is linear!
Exercise
The side-by-side box plot shows a graphical comparison for the two groups. Using this comparison,
answer the following questions.
A. Stock A
B. Stock B
A. Stock A
B. Stock B
A. Stock A
B. Stock B
5. Based on the data in the following table, calculate the Five-number Summary.
Tasks Completed
Population Sample 178 100 153 168 111 91 105 110
a. Minimum:
b. Q1:
c. Median:
d. Q3:
e. Maximum:
7. True or False? A scatterplot always shows the explanatory variable on the horizontal, or x-axis.
a. True
b. False
8. The strength of the correlation between two variables can be measured by which statistic?
a. Relationship factor
b. Correlation coefficient (r)
c. Strength test
d. Trending statistic
Review Checkpoint
To test your understanding of the content presented in this assignment, please click on the
Question icon below. Click your selected response to see feedback displayed below it. If you have
trouble answering, you are always free to return to this or any assignment to re-read the material.
1. A company is studying the effectiveness of two different production processes (Process A and
Process B) for a week. The results are measured by counting the number of products without
defects and with defects. Which numerical measure could be used to analyze the data?
b. Five-number summary
c. Median
d. Conditional percentages
Correct. The answer is d. Both variables are categorical (C → C) so we will use a two-way table.
Therefore, our numerical measure will be conditional percentages.
2. The company also wants to measure the time each process takes. The company measured the
time it took to make each product as it came off of the assembly line. Which numerical measure
would present the maximum and minimum amount of time that it took for products to be made using
each process system?
a. Relative Frequency
b. Five-number Summary
Correct. The answer is b. In this study one variable is categorical, and the other is quantitative
(C → Q). The five-number summary will show five important statistical values: minimum,
maximum, first quartile, median, and third quartile. Therefore, the five-number summary would be
the best choice to look at the minimum and maximum values for both processes.
c. Correlation
d. Conditional percentages
3. A marketer was interested in how age affects responses to images of healthy food. He sampled
a population ranging in age from 10 years old to 70 years old. Subjects were asked to complete a
simple survey, rating how much they would like to eat the food shown in 10 images. The rating went
from 1 (not at all) to 10 (very much). The food depicted was all food considered healthy. Which
numerical measure could be used to analyze the strength of the relationship between age and
response to images of healthy food?
Correct. The answer is a. In this study both variables are quantitative(Q → Q) that form paired
data. The strength of a correlation can be measured by calculating the correlation coefficient (r).
c. Mean
d. Five-number Summary
4. A car company is comparing the miles its trucks can go before needing repair to that of one of its
rivals. The researchers want to compare the middle 50% of the data for both groups. Which
numerical measure would identify the two values that 50% of the data falls between for both
groups?
a. Median
b. Five-number summary
Correct. The answer is b. In this study one variable is categorical and the other is quantitative
(C → Q). The five-number summary will show five important statistical values: minimum,
maximum, first quartile, median, and third quartile. 50% of the data falls between the first(Q1) and
third(Q3) quartile, so the researchers should look at the five-number summary to determine the
values for Q1 and Q3 for both groups, as these values define the middle50% of the data.
c. Mode
d. Joint Frequencies
5. The safety and quality assurance department recently implemented a new product testing
program. The department was interested in determining if this new program decreased a products's
risk of stress related failure within two years. Group A was exposed to the new product testing
program, while the Group B received the traditional testing program. Which numerical measure
could be used to analyze the data?
b. Relative Frequency
c. Joint Frequency
d. Conditional percentages
Correct. The answer is d. In this study both variables are categorical(C → C), so a two-way table
is used to present the data. Therefore, our numerical measure will be conditional percentages.
You may have heard that the speed at which someone adopts new
technology is dependent on his or her age. Said another way, younger
people more readily engage in new technology while older people are less
likely to engage in new technology. There is some truth to this, but the issue
is a lot more complicated. How could you look at data on these two
quantitative variables (age and likelihood of adopting new technology) to
see if there is a relationship? By using a tool called a scatterplot!
Scatterplots are one of the easiest tools we have for comparing two quantitative variables. As you'll
see in this lesson, scatterplots give us a visual depiction of the relationship (or lack of relationship)
between two variables. We use a scatterplot for two quantitative variables because we graph the
relationship on an x-y-plane, similar to what you saw in Module 3. Since we do this on thex-y-
plane, the two variables we're looking at have to be something we can measure, which is exactly
why we use a scatterplot for two quantitative variables.
Other examples of relationships in IT you might want to look at using a scatterplot include:
Remember, if you are looking at two quantitative variables, a scatterplot is a great tool to see how
the two variables compare to one another.
Each pair of (x, y) values appears as a point on the scatterplot, and when we see a completed
scatterplot, it forms a picture of the data that shows us the relationship between the variables. When
we look at the picture, we look for an overall pattern and whether there are any deviations (outliers)
from the pattern. We also can describe the scatterplot by the direction, the form, and the strength of
the relationship.
Positive Correlation
Scatterplot (a) shows a positive correlation between the variables because as thex-
variable increases, the y-variable increases.
Negative Correlation
Scatterplot (c) shows no correlation; there is no apparent overall trend between the two
variables.
Non-Linear Relationship
Direction
Form
If a scatterplot has a pattern of points that form a reasonably straight line, we describe it as
linear. If the points form a pattern that is more curved than straight, we say it is nonlinear,
or curvilinear.
Strength
Here are some examples of scatterplots with descriptions using all three characteristics to
describe the relationship between the x-variable and the y-variable.
Examples
To construct a scatterplot to see any potential relationship between variables, follow the instructions
The data being collected should have two variables. These variables will correspond to thex-axis
and y-axis of your scatterplot. Determine the explanatory variable (a suspected "cause") and the
response variable (a suspected "effect").
For example, a company that specializes in hand-crafted furniture conducts a study to determine a
possible relationship between the size of the bookcases it makes and the average time it takes its
carpenters to assemble the bookcase—that is, the company is investigating if the total board length
(measured in feet) contained in a single bookcase affects the average time (measured in minutes)
needed to assemble that bookcase. Therefore, board length is the explanatory variable on the x-
axis and average assembly time is the response variable on the y-axis. Here is the data that was
collected:
Each dot on the scatterplot represents one style of bookcase, a particular product. The product's
board length and average assembly time can be seen as ordered pairs. The product's board length
measurement is the x-value and the assembly time is the y-value. Here are these points plotted on
a coordinate plane:
Here, that means we can display 15 through 35 on the x-axis (all of these bookcases' board
lengths fall in this range), and we can display 100 through 260 on the y-axis (all of the bookcases'
average assembly times fall in this range). It is also best when creating a scatterplot that the x-axis
and y-axis are both labeled with appropriate units. A scatterplot should have a title over the top of
the graph. Once these changes are made, the scatterplot appears as follows:
There may be more than one outlier, but if several points are grouped together away from the
majority of points, we call them a cluster, not outliers. Consider the following scatterplot:
The presence of an outlier may affect the correlation coefficient r, depending on the placement of
the outlier and the sample size of the data. If an outlier falls far from the regression line, it can
weaken the correlation and move the correlation coefficient r closer to 0. If this outlier is removed,
there will be a stronger relationship between the two variables. On the other hand, if an outlier falls
near the regression line, it will have a diminished effect on the correlation, and may even strengthen
the correlation.
Notice that on the graphs below we have plotted the ordered pairs, as well as a line. This line is
called the "line of best fit" or the "regression line," and it will be discussed in more detail in a later
module. The line of best fit is the line that minimizes the distance between each point and the line
itself. While we can draw many lines on the scatterplot, the line of best fit is the line that is the
closest to all the points in the scatterplot. The closer the points are overall to the line of best fit, the
stronger the correlation (or linear relationship) will be. We are including it here in these graphs to
give you a better feel of the strength of the linear relationship.
An outlier will have an even greater effect if the sample size is smaller. Compare the following two
graphs. First, with the outlier:
5.09 Flashcards
Module 5 Flashcards
Term Definition