Contingency Table and Inference for Two Way Contingency Tables
Categorical Variable
Categorical variables represent types of data which may be divided into groups.
Example
Race, Sex, Age, Group, Educational level, Smoking status etc.
Data
Data is the measurement of characteristics of an object or subject. There are two types of data one
is Qualitative data and another is Quantitative data.
Measurement Scale
Nominal
Categorical variables for which levels do not have a natural ordering are called nominal. For
nominal variables the order of listing of the categories is irrelevant to the statistical analysis.
Example
Religious affiliation (Catholic, Jewish, Protestant, Other), mode of transportation (automobile,
bus, subway, bicycle, other), choice of residence (house, apartment, condominium, other), race,
gender and marital status.
Ordinal
When the categories can be ranked or ordered in an ascending or descending manner+6, such
variables are called ordinal. Ordinal variables clearly order the categories, but absolute distances
between categories are unknown.
Example
Size of automobile (subcompact, compact, mid-size, large), social class (upper, middle, lower),
attitude toward legalization of abortion (strongly disapprove, disapprove, approve, strongly
approve), appraisal of company’s inventory level (too low, about right, too high) etc.
Interval
An interval variable is one that does have numerical distances between any two levels of the scale
and has no true zero.
Example: Temperature, date and time, IQ, CGPA etc.
Ratio
The numerical variable for which there is true zero and the ratio between values is meaningful, is
called ratio level data.
Example: No of patients in a clinic, income, height, weight, distance etc.
Contingency Table
Categorical Response Data ~ 1 of 21
Let denote two categorical response variables, and levels.
When we classify subjects on both variables, there are possible combinations of
classifications. The responses of a subject randomly chosen from some population have a
probability distribution. We display this distribution in a rectangular table having rows for the
categories of and column for the categories of . The cells of the table represent the
possible outcomes. Their probabilities are , where denotes the probability that falls
in the cell in row and column . When the cells contain frequency counts of outcomes, the table
is called a contingency table. The term contingency table was introduced by Karl Pearson (1904).
Another name is cross classification table.
A contingency table having rows and columns is referred to as a by or table.
Y
1 2 … Total
1 …
2 …
X
…
Total …
Where, is the row total, is the column total, is the frequency of the row and
column total.
Notation and Definitions
Let denote the number of observations cross-classified in the cell of the table that is in row
and column and let denote the proportion of the total sample falling in that cell. That is
pij
is the total sample size, so that . is called the joint
probability for ith row and jth column.
pi+ ¿¿
The sample marginal probability for ith row is denoted by , where
p+ j
and the sample marginal probability for jth column is denoted by , where .
Note that and also .
Similar notation will be used for population proportions, with the Greek letter in place of .
π
The joint probabilities for population is denoted by ij and the conditional probabilities for
Categorical Response Data ~ 2 of 21
population is π j ∨i . π j ∨i is the probability that a subject will be classified in jth column of Y when
π ij
it is given that he is already classified in ith row of X. Then, π j ∨i= π ∧∑ π =1 .¿
i+¿ j∨i
j
Table: Notation for joint, conditional and marginal distribution.
Column
1 2 Total
Row 1
2
Total
Independence
When both variables are response variables, we can describe the association using their joint
distribution, the conditional distribution of given , or the conditional distribution of given
.
The conditional distribution of given is related to the joint distribution by for all
and . The variables are statistically independent if all joint probabilities equal the product of
their marginal probabilities that is, if
,
When and are independent,
i.e., each conditional distribution of is identical to the marginal distribution of . Thus, two
variables are independent when the probability of column response is the same in each row, for
. When is a response and is an explanatory variable, the condition
provides a more natural definition of independence than (1).
We use similar notation for sample distributions, with the letter in place of . For instance,
denotes the sample joint distribution in a contingency table. The cell frequencies are denoted
by , with being the total sample size, so
The proportion of times that subjects in row made response is
where .
Categorical Response Data ~ 3 of 21
Methods of Comparing Two Proportions
Difference of proportions
For subjects in row , , is the probability of response 1 and
We use simpler notation π ifor π 1∨i. The difference of proportions of successes, π 1−π 2 ,is a basic
comparison of the two rows.
The difference of proportions falls between -1 to +1. It equals zero when the rows have the
identical conditional distributions. The response Y is independent of the row classification when
π 1−π 2=0
Relative Risk
A difference in proportions of fixed size may have greater importance when both proportions are
close to 0 or 1 than when they are near the middle of the range. For instance, suppose we compare
two drugs in terms of the proportion of subjects who suffer bad side effects. The difference
between 0.010 and 0.001 may be more noteworthy than the difference between 0.410 and 0.401.
in such cases, the ratio of proportions is also a useful descriptive measure.
For tables, the relative risk is the ratio of probabilities.
π1
Realtaive risk =
π2
The ratio can be any non-negative number. A relative risk of 1 corresponding to independence.
.010 10.0∧0.410
For the proportions just given, the relative risks are = =1.02 . Comparison on the
.001 0.401
1−π 1
second response gives a different relative risk, .
1−π 2
Comment
Relative Risk 1 indicates the independency of the categorical variables.
Note:
Relative risk and difference of proportion are affected by the interchange of rows and columns.
Odds Ratio
For a probability π of success, the odds are defined to be
π
odds , Ω=
1−π
The odds are nonnegative, with Ω>1 when a success is more likely than a failure. When π=0.75 ,
0.75
then Ω= =3.0 ,a success is three times as likely as a failure, and we expect about 3
1−0.75
Ω
successes for every one failure. Inversely, ¿ .
Ω+1
πi
In contingence table, within row i, the odds of success instead of failure are Ωi= . The
1−π i
ratio of odds Ω1∧Ω2 in the two rows,
Categorical Response Data ~ 4 of 21
π1
Ω1 1−π 1
θ= =
Ω2 π2
1−π 2
is called odds ratio.
For joint distribution with cell probabilities { π ij } , the equivalent definition for the odds in row i is
πi 1
Ω i= , i=1 ,2. Then the odds ratio is,
πi 2
It is also called the cross-product ratio, since it equals the ratio of the products and of
probabilities from diagonally opposite cells. The odds ratio can equal any non negative number.
When Ω1=Ω2 and , it indicates the independence of X and Y.
When , subjects in row 1 are more likely to have a success than are subjects in row 2; that
is, π 1 >π 2. For instance, when , the odds of success in row 1 are four times the odds in row 2.
This does not mean that the probability π 1=4 π 2, that is the interpretation of a relative risk of 4.0.
When , subjects in row 1 are less likely to have success than in row 2; that is π 1 <π 2.
When one cell has zero probability, .
Values of θ further from 1.0 in a given direction represent stronger association.
Comments
Odds Ratio 1 indicates the independency of the categorical variables.
Properties
The value of does not change if both cell frequencies within any row is multiplies by a
non-zero constant or if both cell frequencies within any column are multiplied by a
constant.
Two values for represent the same level of association, but in opposite directions, when
one value is the inverse of the other. For instance, when , the odds of success in
row 1 are 0.25 times the odds in row 2, or equivalently the odds of success in row 2 are
times the odds in row 1. If the order of the rows is reversed or if the order of
the column is revered, the new value of is simply the inverse of the original value.
It is sometimes more convenient to use . Independence corresponds to . The
log odds ratio is symmetric about this value- reversal of rows or of columns results in a
change in its sign. Two values for that are the same except for sign, such as
and , represent the same level of association.
Categorical Response Data ~ 5 of 21
An implication of the multiplicative invariance property is that the sample odds ratio
estimates the same characteristic even when we select disproportionately large or
small samples from marginal categories of a variable. For instance, suppose a study
investigates the association between vaccination and catching a certain strain of flu. For a
retrospective design, the sample odds ratio estimates the same characteristic whether we
randomly sample (1) 100 people who got the flu and 100 people who did not, or (2) 150
people who got the flu and 50 people who did not, in each case classifying subjects on
whether they took the vaccine. In fact, the odds ratio is equally valid for retrospective,
prospective, or cross-sectional sampling designs. We would estimate the same
characteristic if (3) we randomly sample 100 people who took the vaccine and 100 people
who did not, and then classify them on whether they got the flu, or (4) we randomly
sample 200 people and classify them on whether they took the vaccine and whether they
got the flu.
Example 1:
The following 2×3 contingency table is from a report on the relationship between aspirin use and
heart attacks by the Physicians’ Heath Study Research Group at Harvard Medical School. The
physicians’ Health Study was a 5 year randomized study of whether regular aspirin intake reduces
mortality from cardiovascular disease. Every other day, physicians participating in the study took
either one aspirin tablet or a placebo. The study was blind- those in the study did not know
whether they were taking aspirin or placebo. Of the 11,034 physicians taking a placebo, 18
suffered fatal heart attacks over the course of the study, whereas of the 11,037 taking aspirin, 5
had fatal heart attacks.
Table: Cross-Classification of Aspirin Use and Myocardial Infarction
Aspirin Use Myocardial Infarction Total
Heart Attack No Attack
Placebo 189 10,845 11,034
Aspirin 104 10,933 11,037
Total 293 21,778 22,071
99
189
Proportion of heart attacks among those physicians taking placebo = =0.0171
11,034
104
Proportion of heart attacks among those physicians taking aspirin = =0.0094
11,037
The sample difference in proportions is 0.0171-0.0094 =0.0077
The sample relative risk is 0.0171/0.0094 = 1.82
It indicates that the proportion suffering heart attacks of those taking placebo was 1.82 times the
proportion suffering heart attacks of those taking aspirin.
189 ×10,933
The sample odds ratio is =1.83
104 ×10,845
It indicates that the odds of heart attack for those taking placebo is 1.83 times the odds for those
taking aspirin.
Categorical Response Data ~ 6 of 21
Relationship between Odds Ratio and Relative Risk
π1
Ω1 1−π 1 π 1 (1−π 2) (1−π 2 )
Odds ratio, θ= = = =Relative risk ×
Ω2 π2 π 2 (1−π 1) (1−π 1 )
1−π 2
Their magnitudes are similar whenever the probability π i of the outcome of interest is close to
zero for both groups. We saw this similarity for aspirin study where the heart attack proportion is
less than 0.02 for each group. The relative risk is 1.82 and the odds ratio is 1.83.
When the sampling design is retrospective, it is possible to construct conditional distributions
within levels of the fixed response. It is usually not possible to estimate the probability of the
outcome of interest, or to compute the difference of proportions or relative risk for that outcome.
We can compute the odds ratio, however, since it is determined by the conditional distributions in
either direction. When the probability of the outcome of interest is very small, the population odds
ratio and relative risk take similar values. Thus, we can use the sample odds ratio to provide a
rough indication of the relative risk.
Example 2:
A sample of size 500 respondent was selected in a large metropolitan area to determine various
concerning consumer behavior. The following contingency table was given below:
Enjoys Shopping for clothing Total
Yes No
Sex Male 136 104 240
Female 224 36 260
Total 360 140
a) Find joint probabilities, conditional probabilities.
b) Does shopping depend on sex?
c) What is the probability that a female was not enjoying shopping for clothing?
d) Compute sample difference of proportions, relative risk and odds ratio.
Solution
i) Joint probabilities are,
Categorical Response Data ~ 7 of 21
Also, the conditional probabilities are
ii) Here,
Hence,
So that, we may conclude that enjoy of shopping for clothing depends on sex.
iii) The probability that a female was not enjoy shopping for clothing is
iv) Here, proportion of male is and proportion of female is . So that the
sample difference of proportion is . . So that
proportion enjoy shopping for clothing was 0.658 times lower for male than for female. The
sample odds ratio is .
Measuring Association in I×J Table
Measures of Ordinal Association
A basic question researcher usually poses when analyzing ordinal data is “Does tend to
increase as increases?” Bivariate analyses of interval–scale variables often summarize
Categorical Response Data ~ 8 of 21
covariation by the Pearson correlation, which describes the degree to which has a linear
relationship with . Ordinal variables do not have a defined metric, so the notion of linearity is
not meaningful. However, the inherent ordering of categories allows consideration of monotonic–
for instance, whether tends to increase as does. Measures for ordinal variables that are
analogous to the Pearson correlation describe the degree to which the relationship is monotone.
In a strict sense, comparisons of two subjects on an ordinal scale can answer “Which subject
makes the higher response?” when we observe the ordering of two subjects on each of two
variables, we can classify the pair of subjects as concordant or discordant.
Concordant and Discordant
The pair is concordant if the subject ranking higher on variable also ranks higher on variable
. The pair is discordant if the subject ranking higher on ranks lower on . The pair is tied if the
subjects have the same classification on and / or .
Consider two independent observations from a joint probability distribution for two ordinal
variables. For that pair of observations,
are the numbers of concordance and discordance respectively.
Several measures of association for ordinal variables utilize the difference between these
probabilities. For these measures, the association is said to be positive if and negative
if and independent if .
Example of Job Satisfaction
We illustrate concordance and discordance using Table below, taken from the 1984 General
Social Survey of the National Data Program in the United States as quoted by Norusis (1988). The
variables are income and job satisfaction. Income has levels less than
and and denoted by
, and over . Job satisfaction has levels very dissatisfied (DV), little
dissatisfied (LD), moderately satisfied (MS), and very satisfied (VS). We treat VS as the high end
of the job satisfaction scale.
Table: Cross Classification of Job Satisfaction by Income
Job Satisfaction
Very Little Moderately Very
Income Dissatisfied Dissatisfied Satisfied Satisfied
20 24 80 82
22 38 104 125
Categorical Response Data ~ 9 of 21
13 28 81 113
7 18 54 92
Consider a pair of subjects, one of whom is classified in the cell and the other in the cell
, so there are concordant pairs from these two cells. The 20 subjects in the
cell are also part of a concordant pair when matched with each of the other
subjects ranked higher on both variables. Similarly, the 24
subjects in the cell cell are part of concordant pairs when matched with the
subjects ranked higher on both variables.
The total number of concordant pair denoted by
The number of discordant pairs of observations is
In this example, suggests a tendency for low income to occur with low job satisfaction and
high income with high job satisfaction.
Gamma
Given that the pair is untied on both variables, is the probability of concordance and
is the probability of discordance. The difference between these probabilities is
called Gamma. Its range is . The absolute value of the correlation is 1 when the
relationship between and is perfectly linear, only monotonicity is required for , with
, if and if . The perfect association value occurs even when the
relationship is not strictly monotone. If , for instance, then for observations and
on a pair of subjects and having , it follows that but not necessarily
that . Independence implies , but the converse is not true.
Yule’s Coefficient
Categorical Response Data ~ 10 of 21
For tables, we define . This measure, which Yule (1900, 1912) introduced
and called in honor of the Belgian statistician Quetelet, is now referred to as Yule’s . the
range of is .
Kendall’s tau-b and Somers’d
The sample correlation between the distinct pairs equals
This index of ordinal association is called Kendall’s tau-b. Tau-b tends to be less sensitive than
gamma to the choice of response categories.
So,
Hence, Somers’d can be also defined as
Where, indicate the difference between the proportion of concordant and discordant pairs, out
of those pairs untied on .
Distributions for Categorical Data
Inferential data analysis require assumptions about the random mechanism that generated the data. Three
distributions that are mostly used with categorical data are binomial, Poisson and multinomial
distribution.
Sampling Distribution
Categorical Response Data ~ 11 of 21
Suppose we observe counts in the cells of a contingency table. For instance,
these might be observations for the levels of a single categorical variable, or for cells
of a two-way table. Let the counts be random variables. Each has distribution concentrated on
the nonnegative integers, with expected value denoted by . The are called expected
frequencies.
Poisson Sampling
Since must be a non-negative integer, its sampling distribution should place its mass on that
range. One of the simplest such distributions is the Poisson. Its form depends on a single
parameter, the mean . The probability mass function is
It satisfies .
The Poisson sampling model for counts assumes that they are independent Poisson random
variables. The joint probability function for is then the product of the probabilities for the
cells. The total sample size also has a Poisson distribution with parameter .
The Poisson distribution is used for counts of events that occur randomly over time or space,
when outcomes in disjoint periods are independent. For example, Poisson distributions might be
realistic for
The number of spontaneous abortions,
The number of induced abortions,
The number of live births measured in January, 2005 in Dhaka.
Multinomial Sampling
An unusual feature of Poisson sampling is that the total sample size is ransom, rather
than fixes. If we start with the Poisson model but condition on the total sample size , no
longer have Poisson distributions, since each are also no longer independent, since the value of
one affects the possible range for the others.
Given that , the conditional probability of a set satisfying this condition is
Categorical Response Data ~ 12 of 21
This is the multinomial distribution, characterized by the sample size and the cell
probabilities . The Binomial distribution with index and “success” probability is
the special case of the multinomial with cells. For the multinomial distribution for
, the marginal distribution for is Binomial, with and .
The multinomial distribution for also applies when independent observations are taken
from a probability distribution concentrated on a set of categories. In other words, if the same
probability distribution applies to each observation, and if the observations are
independent, then the counts of the number of observations in each category have
distribution . When cell counts have distribution , the sampling scheme is calls multinomial
sampling.
Confidence Intervals for Association Parameters
Estimating Odds Ratios and its Confidence Interval
Let denote the sample value of the odds ratio for a table. The sample
odds ratio equals if any and it is undefined if both entries in a row or column are
zero.
In terms of bias and mean squared error, Gart and Zweiful (1967) and Haddane (1955) showed
that
behaves well. The log transform, having an additive rather than multiplicative structure,
converges more rapidly to a normal distribution. For Poisson or multinomial sampling or for
independent Binomial sampling within the rows or within the columns, an estimated asymptotic
standard error (ASE) of is
By the large sample normality of log ( θ^ ) ,the Wald confidence interval for log(θ ¿ is
Categorical Response Data ~ 13 of 21
^ z α σ^ ¿
log ( θ)± )
2
Estimating Difference of Proportions and Relative Risk
Let us consider the difference of proportions and the relative risk for comparing conditional
distributions of a column response variable within two rows. For these measures, we treat the
rows as independent Binomial samples. For group i, Y i has a binomial distribution with sample
size ni and a probability π i of a success outcome.
yi
The sample proportion ^π i= has expectation π i and variance π i ¿ )/ni .Since ^π 1 and ^π 2 are
ni
independent, their difference has E( π^ ¿ ¿ 1−π^ 2)=π 1−π 2 ¿and standard error
σ ( π^ 1− π^ 2 ) =
√ π 1 (1−π 1 ) π 2 (1−π 2)
n1
+
n2
The estimate σ^ ( π^ 1− π^ 2 ) replaces π i by π^ i . Then 100(1-α) percent confidence interval for
π 1−π 2 is
( π^ 1− π^ 2 ) ± z α σ^ ( ^π 1−^π 2 )
2
y1
π^ 1 n1
The sample relative risk is ¿ = . Like the odds ratio, it converges to normality faster on the
π^ 2 y 2
n2
log scale. An estimated standard error for log r is
σ^ ¿
^
The Wald interval exponentials endpoints of log r ± z α σ ¿.
2
Example 1 (continued):
Aspirin Use Myocardial Infarction Total
Fatal Attack Non-fatal No Attack
Attack
Placebo 18 ( y 1 ¿ 171 10,845 11,034 (n1 ¿
Aspirin 5 ( y2 ) 99 10,933 11,037 (n2 ¿
Total 23 270 21,778 22,071
From the previous table 1 of aspirin use and heart attacks, we get,
18
The proportion having fatal heart attacks for those taking placebo is ^π 1= =0.00163
11,034
5
The proportion having fatal heart attacks for those taking aspirin is ^π 2= =0.00045
11,037
π^ 1 0.00163
The sample relative risk is r = = =3.6
π^ 2 0.00045
The 95% confidence interval for the log relative risk using σ^ ¿ =0.505 is
log (3.6)±1.96 (0.505)
Categorical Response Data ~ 14 of 21
This translates to (1.34, 9.70) for the relative risk. This indicates that the death rate for those
taking placebo is between 1.34 and 9.70 times that for those taking aspirin.
Again,
( π^ 1− π^ 2 ) =0.00163−0.00045=0.0012
σ ( π^ 1− π^ 2 ) =
√ n1
+
n2 √
π 1 ( 1−π 1 ) π 2 ( 1−π 2 )
=
0.00163 ( 1−0.00163 ) 0.00045 ( 1−0.00045 ) =0.00043
11034
+
11037
The Wald 95% CI for π 1−π 2 is 0.0012 ± 1.96 (0.00043) = (0.0003, 0.0020)
The relative risk is more useful than π 1−π 2 for these data, because the rates of heart attack death
is very low but with ratio quite far from 1.
Testing Goodness of fit
Testing a Specified Multinomial
A goodness of fit test introduced by Karl-Pearson in 1900, Consider the null hypothesis, that the
parameter of a multinomial distribution equal certain fixed values , where
. When is true, the expected cell frequencies are . For
sample counts , Pearson proposed the test statistic
.
For large samples, has approximately a chi-squared null distribution with degrees of freedom
equal to . A statistic of form is called a Pearson chi-squared statistic.
Testing Independence in Two Way Contingency Table
Pearson chi-squared test
We estimate the expected frequencies by . The statistic then equals
. Pearson (1900, 1922) claimed that replacing by the estimates
would not affect the distribution of . Since there are categories for the cross-
classification, he argued that would have an asymptotic chi-squared distribution with
. On the contrary, since are determined by estimating and , the chi-
squared distribution has .
Likelihood Ratio chi-squared
Categorical Response Data ~ 15 of 21
The likelihood ratio test is a general purpose way of testing a null hypothesis against an
alternative hypothesis . In this test, we maximize the likelihood under is true. Let
denote the ratio of the maximized likelihoods, which cannot exceed 1. Wilks showed
that has a limiting null chi-squared distribution as .
For multinomial sampling in a contingency table, the kernel of the likelihood is , where
all and
Under independence , the likelihood is maximized when and
, so that . In the general case, the likelihood is maximized when . So
that we get,
Again,
Thus the ratio of the likelihoods equals
It follows that Wilk’s statistic, denoted by is
Categorical Response Data ~ 16 of 21
This statistic is called the likelihood ratio chi-squared statistic. The larger value of , the mire
evidence there is against the null hypothesis. For large samples, has a chi-squared null
distribution with
Table : Attained education and belief in God
Highest Belief in God Total
academic Don’t No way Some Believe Believe but Know
degree belief to find higher sometimes doubts God
out power exists
Less than 9 8 27 8 47 236 n1 +¿=335 ¿
High School
High School 23 39 88 49 179 706 n2 +¿=1084 ¿
or Junior
College
Graduate n3 +¿=581 ¿
Total n+1 =60 n+2 =95 n+3 =204 n+ 4=76 n+5 =¿330 n+6 =1235 n =2000
Here, the null hupothesis is
H o :There is no association between education and belief in God, hence, independence
H 1 : There is association between education and belief in God
n1 +¿n 335∗60
^ 11=
Observed value for (1,1) cell is, m +1
= =10 ¿
n 2000
And so on.
=76.1
2
And G = 73.2 with df=(3-1)(6-1)=10
Critical values, χ 20.05 , 10=18.31
As, Calculated value> Tabulated value, so the null hypothesis is rejected.
So it may be concluded that there exists an association between education level and belief in God.
In SPSS, first go to variable view.
Variable view
Categorical Response Data ~ 17 of 21
Then go to data view and input the data.
Data view
Categorical Response Data ~ 18 of 21
Categorical Response Data ~ 19 of 21
Categorical Response Data ~ 20 of 21
Categorical Response Data ~ 21 of 21