Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
UNIT IV BIVARIATE ANALYSIS
Relationships between two variables - percentage tables - analyzing contingency tables -
handling several batches - scatterplots and resistant lines – transformations.
RELATIONSHIPS BETWEEN TWO VARIABLES
Correlation refers to a process for establishing the relationships between two
variables. You learned a way to get a general idea about whether or not two
variables are related, is to plot them on a “scatter plot”. While there are many
measures of association for variables which are measured at the ordinal or higher
level of measurement, correlation is the most commonly used approach.
Methods of correlation summarize the relationship between two variables in a
single number called the correlation coefficient. The correlation coefficient is
usually represented using the symbol r, and it ranges from -1 to +1.
A correlation coefficient quite close to 0, but either positive or negative, implies
little or no relationship between the two variables. A correlation coefficient
close to plus 1 means a positive relationship between the two variables, with
increases in one of the variables being associated with increases in the other
variable.
A correlation coefficient close to -1 indicates a negative relationship between
two variables, with an increase in one of the variables being associated with a
decrease in the other variable. A correlation coefficient can be produced for
ordinal, interval or ratio level variables, but has little meaning for variables
which are measured on a scale which is no more than nominal.
For ordinal scales, the correlation coefficient can be calculated by using
Spearman’s rho. For interval or ratio level scales, the most commonly used
correlation coefficient is Pearson’s r, ordinarily referred to as simply the
correlation coefficient.
Types of Correlation
The scatter plot explains the correlation between the two attributes or variables.
It represents how closely the two variables are connected. There can be three
such situations to see the relation between the two variables –
• Positive Correlation – when the values of the two variables move in the
same direction so that an increase/decrease in the value of one variable
is followed by an increase/decrease in the value of the other variable.
• Negative Correlation – when the values of the two variables move in
the opposite direction so that an increase/decrease in the value of one
variable is followed by decrease/increase in the value of the other
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
variable.
• No Correlation – when there is no linear dependence or no relation
between the two variables.
A bivariate table displays the distribution of one variable across the categories of
another variable. It is obtained by classifying cases based on their joint scores
on two nominal or ordinal variables. It can be thought of as a series of frequency
distributions joined to make one table. The data in Table 9.1 represent a sample
of General Social Survey (GSS) respondents by race and whether they own or rent
their homes (in this case, both variables are nominal-level measurements). To
make sense of these data, we must first construct the table in which these
individual scores will be classified. In Table 9.2, the 17 respondents have been
classified according to joint scores on race and home ownership. The table has
the following features typical of most bivariate tables: 1. The table’s title is
descriptive, identifying its content in terms of the two variables.
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
2. It has two dimensions, one for race and one for home ownership. The variable
home ownership is represented in the rows of the table, with one row for owners
and another for renters. The variable race makes up the columns of the table,
with one column for each racial group. A table may have more columns and
more rows, depending on how many categories the variables represent. For
example, had we included a group of Latinos, there would have been three
columns (not including the row total column). Usually, the independent variable
is the column variable and the dependent variable is the row variable.
3. The intersection of a row and a column is called a cell. For example, the
two individuals represented in the upper left cell are blacks who are also
home owners.
4. The column and row totals are the frequency distribution for each variable,
respectively. The column total is the frequency distribution for race, the row
total for home ownership. Row and column totals are sometimes called
marginals. The total number of cases (N) is the number reported at the
intersection of the row and column totals. (These elements are all labelled
in the table.)
5. The table is a 2 × 2 table because it has two rows and two columns (not counting
the marginals). We usually refer to this as an r × c table, in which r represents the
number of rows and c the number of columns. Thus, a table in which the row
variable has three categories and the column variable has two categories would
be designated as a 3 × 2 table.
Bivariate table:
A table that displays the distribution of one variable across the categories of
another variable.
Column variable:
A variable whose categories are the columns of a bivariate table.
Row variable:
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
A variable whose categories are the rows of a bivariate table.
Cell:
The intersection of a row and a column in a bivariate table.
Marginals:
The row and column totals in a bivariate table
HOW TO COMPUTE PERCENTAGES IN A BIVARIATE TABLE
To compare home ownership status for blacks and whites, we need to convert
the raw frequencies to percentages because the column totals are not equal.
Recall from Chapter 2 that percentages are especially useful for comparing two
or more groups that differ in size.
There are two basic rules for computing and analysing percentages in a
bivariate table:
1. Calculate percentages within each category of the independent variable.
2. Interpret the table by comparing the percentage point difference for
different categories of the independent variable.
Calculating Percentages Within Each Category of the Independent Variable
The first rule means that we have to calculate percentages within each
category of the variable that the investigator defines as the independent
variable. When the independent variable is arrayed in the columns, we
compute percentages within each column separately. The frequencies within
each cell and the row marginals are divided by the total of the column in which
they are located, and the column totals should sum to 100%. When the
independent variable is arrayed in the rows, we compute percentages within
each row separately. The frequencies within each cell and the column
marginals are divided by the total of the row in which they are located, and the
row totals should sum to 100%.
In our example, we are interested in race as the independent variable and in
its relationship with home ownership. Therefore, we are going to calculate
percentages by using the column total of each racial group as the base of the
percentage. The percentage of black respondents who own their homes is
obtained by dividing the number of black home owners by the total number
of blacks in the sample.
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
Table 9.3 presents percentages based on the data in Table 9.2. Notice that the
percentages in each column add up to 100%, including the total column
percentages. Always show the Ns that are used to compute the percentages—in
this case, the column totals.
Comparing the Percentages Across Different Categories of the Independent
Variable
The second rule tells us to compare how home ownership varies between
blacks and whites. Comparisons are made by examining differences between
percentage points across different categories of the independent variable.
Some researchers limit their comparisons to categories with at least a 10-
percentage point difference. In our comparison, we can see that there is a
41.4 percentage point difference between the percentages of white home
owners (70%) and black home owners (28.6%). In other words, in this group,
whites are more likely to be home owners than blacks.4 Therefore, we can
conclude that one’s race appears to be associated with the likelihood of being
a home owner.
Note that the same conclusion would be drawn had we compared the
percentage of black and white renters. However, since the percentages of
home owners and renters within each racial group sum to 100%, we need to
make only one comparison. In fact, for any 2 × 2 table, only one comparison
needs to be made to interpret the table. For a larger table, more than one
comparison can be made and used in interpretation
HOW TO DEAL WITH AMBIGUOUS RELATIONSHIPS BETWEEN VARIABLES
Sometimes it isn’t apparent which variable is independent or dependent;
sometimes the data can be viewed either way. In this case, you might compute
both row and column percentages. For example, Table 9.4 presents three sets
of figures for the variables SPANKING and FEFAM for a sample of 127 GSS
respondents: (a) the absolute frequencies, (b) the column percentages,
and (c) the row percentages. SPANKING is measured with the survey question
“Do you favour spanking to discipline a child?” The variable FEFAM measures
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
whether the respondent agrees or disagrees with the statement “a man should
work and a woman should stay at home.” Table
9.4b shows that respondents who strongly disagree with spanking a child are
less likely to agree with the FEFAM statement than those who strongly agree with
spanking (10% compared with 50%). Table 9.4c shows that individuals who
strongly agree that a man should work and a woman should stay at home are
more likely to agree with spanking than those who disagree with the statement
on men’s and women’s roles (94% compared with 63%).
Thus, percent aging within each column (Table 9.4b) allows us to examine the
hypothesis that spanking (the independent variable) is associated with
agreement with the FEFAM statement (the dependent variable). When we
percentage within each row (Table 9.4c), the hypothesis is that agreement or
disagreement with the FEFAM statement (the independent variable) may be
related to SPANKING (the dependent variable).5
Finally, it is important to understand that ultimately what guides the
construction and interpretation of bivariate tables is the theoretical question
posed by the researcher. Although the particular example in Table 9.4 makes
sense if interpreted using row or column percentages, not all data can be
interpreted this way. For example, a table comparing women’s and men’s
attitudes toward sexual harassment in the workplace could provide a sensible
explanation in only one direction. Gender might influence a person’s attitude
toward sexual harassment; however, a person’s attitude toward sexual
harassment certainly couldn’t influence her or his gender. Therefore, either row
or column percentages are appropriate, depending on the way the variables
are arrayed, but not both.
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
What is a Contingency Table?
A contingency table displays frequencies for combinations of two categorical
variables. Analysts also refer to contingency tables as crosstabulation and two-
way tables.
Contingency tables classify outcomes for one variable in rows and the other in
columns. The values at the row and column intersections are frequencies for
each unique combination of the two variables.
Use contingency tables to understand the relationship between categorical
variables. For example, is there a relationship between gender (male/female)
and type of computer (Mac/PC)?
I love these tables because they organize your data and allow you to answer
diverse questions. In this post, learn about contingency tables, including how to
interpret, graph, and analyse them.
Example Contingency Table
The contingency table example below displays computer sales at our fictional
store. Specifically, it describes sales frequencies by the customer’s gender and
the type of computer purchased. It is a two-way table (2 X 2). I cover the naming
conventions at the end.
In this contingency table, columns represent computer types and rows represent
genders. Cell values are frequencies for each combination of gender and
computer type. Totals are in the margins. Notice the grand total in the bottom-
right margin.
At a glance, it’s easy to see how two-way tables both organize your data and paint
a picture of the results. You can easily see the frequencies for all possible subset
combinations along with totals for males, females, PCs, and Macs.
For example, 66 males bought PCs while females bought 87 Macs. Furthermore,
there are 117 females, 106 males, 96 PC sales, 127 Mac sales, and a grand total
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
of 223 observations in the study.
Marginal and Conditional Distributions in Contingency Tables
Contingency tables are a fantastic way of finding marginal and conditional
distributions. These two distributions are types of frequency distributions.
Learn more about Frequency Tables: How to Make and Interpret.
Marginal Distribution
These distributions represent the frequency distribution of one categorical
variable without regard for other variables. Unsurprisingly, you can find these
distributions in the margiws of a contingency table.
The following marginal distribution examples correspond to the blue highlights.
For example, the marginal distribution of gender without considering computer
type is the following:
o Males: 106
o Females: 117
Alternatively, the marginal distribution of computer types is the following:
o PC: 96
o Mac: 127
Learn more about Marginal Distributions.
Conditional Distribution
For these distributions, you specify the value for one of the variables in the
contingency table and then assess the distribution of frequencies for the other
variable. In other words,
you cowditiow the frequency distribution for one variable by setting a value of
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
the other variable. That might sound complicated, but it’s easy using a
contingency table. Just look across one row or down one column.
The following conditional distribution examples correspond to the green
highlights.
For example, the conditional distribution of computer type for females is the
following:
o PC: 30
o Mac: 87
Alternatively, the conditional distribution of gender for Macs is the following:
o Males: 40
o Females: 87
Learn more about Conditional Distributions.
Finding Relationships in a Contingency Table
In the contingency table below, the two categorical variables are gender and ice
cream flavour preference. This is a two-way table (2 X 3) where each cell
represents the number of times males and females prefer a particular ice
cream flavour. The CSV datasheet shows one format you can use to enter the
data into your software: Flavour Preference.
How do we go about identifying a relationship between gender and flavour
preference?
If there is a relationship between ice cream preference and gender, we’d expect
the conditional distribution of flavours in the two gender rows to differ. From
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
the contingency table, females are more likely to prefer chocolate (37 vs. 21),
while males prefer vanilla (32 vs. 12). Both genders have an equal preference
for strawberry. Overall, the two-way table suggests that males and females
have different ice cream preferences.
The Total column indicates the researchers surveyed 66 females and 71 males.
Because we have roughly equal numbers, we can compare the raw counts
directly. However, when you have unequal groups, use percentages to
compare them.
Row and Column Percentages in Contingency Tables
Row and column percentages help you draw conclusions when you have
unequal numbers in the margins. In the contingency table example above,
more women than men prefer chocolate, but how do we know that’s not due
to the sample having more women? Use percentages to adjust for unequal
group sizes. Percentages are relative frequencies. Learn more about Relative
Frequencies and their Distributions.
Here’s how to calculate row and column percentages in a two-way table.
o Row Percentage: Take a cell value and divide by the cell’s row total.
o Column Percentage: Take a cell value and divide by the cell’s column
total.
For example, the row percentage of females who prefer chocolate is simply the
number of observations in the Female/Chocolate cell divided by the row total for
women: 37 / 66 = 56%.
The column percentage for the same cell is the frequency of the
Female/Chocolate cell divided by the column total for chocolate: 37 / 58 =
63.8%.
Interpreting Percentages in a Contingency Table
The contingency table below uses the same raw data as the previous table and
displays both row and column percentages. Note how the row percentages sum
to 100% in the right margin while the column percentages sum to 100% at the
bottom.
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
Whether you focus on row percentages or column percentages in a contingency
table depends on the question you’re answering. In our case, we want to know
whether flavour preference depends on gender. Because the two genders
display in separate rows, we’ll look for differences in the row percentages.
56% of females prefer chocolate versus only 29.6% of males. Conversely, 45%
of males prefer vanilla, while only 18.2% of females prefer it. These results
reconfirm our previous findings using the raw counts.
How to Graph a Contingency Table
You can use bar charts to display a contingency table. The following clustered
bar chart shows the row percentages for the previous two-way table. I’ve set
the graph to cluster the female and male pairs of bars together for each flavour,
making comparisons easier. I think it gives a nice oomph to the tabular results.
This bar chart reiterates our conclusions from the contingency table. Women in
this sample prefer chocolate, men favour vanilla, and both genders have an
equal preference for strawberry.
How to Analyse a Contingency Table
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
We’ve already looked at various ways to analyse a contingency table. Here are
two more methods that take it to another level.
Contingency tables are a fantastic way to display and find various types of
probabilities. Use these tables to calculate joint, marginal, and conditional
probabilities. I’ve written an article about calculating probabilities using two-way
tables, and it includes all the definitions, notation, and formulas you need. Read
about Using Contingency Tables to Calculate Probabilities.
In this post, we looked for a relationship between gender and ice cream
preference by noting the differences between counts and row percentages in
the contingency table. If we’re using this sample to draw inferences about the
entire population of ice cream consumers, we’ll need to use a hypothesis test
to evaluate the relationship.
In other words, are the differences we noticed in the sample large enough to
support the notion that a relationship exists in the population? Or can we chalk
up the differences
to random sampling error? Learn how the chi-square test of independence can
help us out by analysing contingency tables!
Naming Conventions for Contingency Tables
Contingency tables come in a variety of flavours. The key considerations for
naming the types are the number of categorical variables and the number of
values for each categorical variable.
Number of Categorical Variables
You must have at least two categorical variables to create a contingency table.
When you have
two variables, it’s a two-way table. If you have three, it’s a three-way table, and
so on.
For example, suppose we run a computer store and record the sales using the
two categorical variables of gender and computer type. Those variables create
a two-way contingency table. If we add a third categorical variable for store
location, it becomes a three-way table.
How do you present a three-way table?
Selvam College of Technology (Autonomous),
“A” Grade by NAAC, UGC recognized 2(f) Status, Approved by AICTE – New Delhi, Affiliated to Anna University
Namakkal – 03. [Link]
Because contingency tables display in two dimensions, you need multiple tables
to represent anything more than a two-way table.
If there are four store locations in our three-way example, we’ll need to use four
two-way tables. Each table displays gender and computer type for one
location.
Number of Rows and Columns
The rows represent values of one categorical variable, while the columns
denote the values of another. Analysts indicate the number of values for each
variable by describing these tables as an A X B contingency table, where A
represents the number of rows and B signifies the number of columns.
For example, at our computer store, Gender has two possible values and
computer type has two values (PC and Mac). Hence, we have a 2 X 2
contingency table. If we add a third type of computer as a new column, it
becomes a 2 X 3 table.