Introduction
Categorical data analysis
Agnes Tuti R
• Categorical (or discrete)
variables are used to organize
observations into groups that
share a common trait.
Description Categorical data analysis is the
analysis of data where the
• The trait may be nominal (e.g.,
sex or eye color) or ordinal (e.g.,
response variable has been age group), and, in general, the
grouped into a set of mutually number of groups within a
exclusive ordered (such as age variable is 20 or fewer (Imrey &
group) or unordered (such as Koch, 2005).
eye color) categories. • Most statistical procedures
distinguish between
independent, or explanatory,
and dependent, or response,
variables. For instance,
an analysis of variancemay be
used to determine how a
continuous response variable
varies according to explanatory
variable levels such as eye color.
In contrast, categorical data
analysis involves the statistical
treatment of categorical
response variables.
Measurement Scale
• 1. Nominal
Qualitative Data
• 2. Ordinal
• 3. Interval
Quantitaive Data
• 4. Ratio
What is Categorical Data?
Types of Categorical Data
Nominal Data
Nominal data
Types of Categorical Data
2. Ordinal Data
it is said to exhibit both
categorical and numerical data characteristics
General Characteristics/Features of Categorical Data
• Categories
These consist of two categories of categorical data, namely; nominal data and ordinal data.
Nominal data, also known as named data is the type of data used to name variables,
while ordinal data is a type of data with a scale or order to it.
• Qualitativeness
Categorical data is qualitative. That is, it describes an event using a string of words
rather than numbers.
• Analysis
Categorical data is analysed using mode and median distributions, where nominal data is
analysed with mode while ordinal data uses both. In some cases, ordinal data may also be
analysed using univariate statistics, bivariate statistics, regression applications, linear
trends and classification methods.
• Graphical analysis
It can also be analysed graphically using a bar chart and pie chart. A bar chart is mostly
used to analyse frequency while a pie chart analysis percentage. This is done after
grouping it into a table.
General Characteristics/Features of Categorical Data
• Interval scale
In the case of ordinal data, which has a given order or scale, the scale does not have a
standardised interval. This is not applicable for nominal data.
• Numeric values
Although categorical data is qualitative, it may sometimes take numerical values.
However, these values do not exhibit quantitative characteristics. Arithmetic operations
can not be performed on them.
• Nature
Categorical data may also be classified into binary and non-binary depending on its nature.
A given question with options “Yes” or “No” is classified as binary because it has
two options while adding “Maybe” to the given options will make it non-binary.
1. Household Income:
Categorical data is mostly
used by businesses when
investigating the spending
power of their target
audience, to conclude on an
Categorical affordable price for their
Data products. For example:
Examples
2. Education Level: The level
of education of a respondent
may be requested for when
filling forms for job
applications, admission,
training etc. This is used to
assess their qualification for a
specific role. Consider the
example below:
3. Gender:
Respondents are asked for their gender when
filling out a biodata. This is mostly categorised as
male or female, What is your gender?
• Male
• Female
This is a binary and closed-ended nominal data
example but may also be nonbinary. For
example:
4. Customer satisfaction:
After rendering service to customers, businesses
like to get feedback from customers regarding
their service to improve. For example;
5. Brand of soaps:
When doing competitive analysis
research, a soap brand may want to study
the popularity of its competitors among
its target audience. In this case, we have
something of this nature:
Syarat PENGKATEGORIAN SUATU VARIABEL
1. HOMOGEN
Categories in one variable are the same object.
2. MUTUALLY EXCLUSIVE
among categories are mutually exclusive
3. MUTUALLY EXHAUSTIVE
Complete decomposition down to the smallest unit
4. SKALA NOMINAL ATAU ORDINAL
1< X<5
5 < X < 10
CONTINGENCY TABLES
Dependency of two variables
Cross classifies a sample of Indonesians according to their gender and their opinion about the
pandemic of Corona 19 in Indonesia has ended and has entered an endemic period. For the females
in the sample, for example, 509 said they believed that know Indonesia is entered an endemic
period and 116 said they did not believe or were undecided. Does an association exist between
gender and belief in an endemic? Is one gender more likely than the other to believe in an endemic,
or is belief in an endemic independent of gender?
Belief in Endemic Period
Gender Total
Yes No or Undecided
Male 509 116 625
Female 398 104 502
Total 210 1127
PROBABILITY STRUCTURE FOR CONTINGENCY TABLES
• Suppose there are two categorical variables, denoted by X and Y . Let I
denote the number of categories of X and J the number of categories of Y .
• A rectangular table having I rows for the categories of X and J columns for
the categories of Y has cells that display the I J possible combinations of
outcomes.
• A table of this form that displays counts of outcomes in the cells is called a
contingency table.
• A table that cross classifies two variables is called a two-way contingency
table; one that cross classifies three variables is called a three-way con-
tingency table, and so forth.
• A two-way table with I rows and J columns is called an I × J (read I–by–J)
table. Table 1 is a 2 × 2 table.
Joint, Marginal, and Conditional Probabilities
• Probabilities for contingency tables can be of three types:
• joint,
• marginal, or
• condi- tional.
• Suppose first that a randomly chosen subject from the population of interest is classified
on X and Y.
• Let πij = P(X = i,Y = j) denote the probability that (X, Y ) falls in the cell in row i and
column j .
• The probabilities {πij } form the joint distribution of X and Y . They satisfy ∑ 𝜋!" = 1.
• The marginal distributions are the row and column totals of the joint probabilities. We
denote these by {πi+ } for the row variable and {π+j} for the column variable, where the
subscript “+” denotes the sum over the index it replaces. For 2 × 2 tables,
π1+ =π11 +π12 and π+1 =π11 +π21
Each marginal distribution refers to a single variable.
I X J CONTINGENCY TABLES (population)
X
Total
Y KATAGORI (J)
1 2 3 . . j
K 1 π11 π12 π13 . . π1j Π1+
A
T 2 π21 π22 π23 Π2+
A (I)
G
.
O i πij Πi+
R
.
Y
TOTAL Π+1 Π+2 Π+3 Π+j Π++
• We use similar notation for samples, with p in place of π . For example, {pij } are
cell proportions in a sample joint distribution.
• We denote the cell counts by {nij}. The marginal frequencies are the row totals
{ni+} and the column totals relate to the cell counts by pij =nij/n
• In many contingency tables, one variable (say, the column variable, Y ) is a
response variable and the other (the row variable, X) is an explanatory variable.
• Then, it is informative to construct a separate probability distribution for Y at
each level of X. Such a distribution consists of conditional probabilities for Y ,
given the level of X. It is called a conditional distribution.
jurusan Statistika S1 ITS 13/02/2019
njutan Contingency
Tabel KontingensiTable
18
... 18 jurusan Statistika S1 ITS
jurusan Statistika S1 ITS 13/02/2019
• Two Dimensional (r x c) Tabel Probabilitas 2 DIMENSI (2 x 2)
jurusan Statistika S1 ITS 13/02/2019 jurusan Statistika S1 ITS 13/0
Tabel KONTINGENSI
Tabel• Three
KONTINGENSI 2 DIMENSI
2 DIMENSI (2 x 2)
Dimensional (r x cxl) (2 xTabel
2) Probabilitas 2 DIMENSI (2 x 2)
da 2 variabel : variabel A dan B
There are 4 possible even:
Example : 2x2 contingency table B1 A2B2
A1B1, A1B2, A2B1 dan B2 Total
B1 B2 Total
ariabel A mempunyai
B1 2 Total : A1 danAA
B2kategori 1 2 B1 p11 B2 p12 p1+
Total
A1 n11 n12 n1+
ariabel B
A1 mempunyai
A2
n11
n21
2
n12kategori
n22
n1+ : B dan
n2+ 1
A1 AB
2
2
p11 p21 p12p22 p1+p2+
A2 n21 n22 n2+ A2Total p21 p+1 p22p+2 p2+p++
Total n+1 n+2 n++
ehingga kemungkinan yang
Total n+1 n +2
terjadi
n ++
: Total p+1 p+2 p++
A1B1, A1B2, A2B1 dan A2B2
2 2 JikaIfA dan
A and B areBindependent,
saling bebas,
so maka :
n1+ =2 ∑ n1 j
2 2
n+1 = ∑ ni1
2
n++ =2 ∑∑ nij
n1+ = ∑ jn=11 j
2
j =1
n+1 = ∑in=1i1 n++ = ∑∑ i =1 jn
=1ij Jika A dan p B= saling
p p bebas, maka :
i =1 i+ +j
nij
i =1 j =1 ij
robabilitas tiapofsel
Probability each :Cell = pij = pij = pi + p+ j
n++
I X J CONTINGENCY TABLES (sample)
X
Total
Y KATAGORI (J)
1 2 3 . . j
K 1 p11 p12 p13 . . p1j p1+
A
T 2 p21 p22 p23 p2+
A (I)
G
.
O i pij pi+
R
.
Y
TOTAL p+1 p+2 p+3 p+j p++
Example
Belief in Endemic (Y) Belief in Endemic (Y)
Gender Total Gender Total
(X) (X)
Yes No or Undecided Yes No or Undecided
Male 509 116 625 Male 0.452 0.103 0.555
Female 398 104 502 Female 0.353 0.092 0.455
Total 907 210 1127 Total 0.805 0.195 1
Marginal Distribution
Belief in Endemic Gender
0,555
0,805 0,455
0,195
yes no or undecided MALE FEMALE
Conditional Distribution
• Biliefeness in Endemic for female
Belief in Endemic (Y)
Gender Total Belief in Endemic (Y)
(X) Gender Total
Yes No or Undecided
(X)
Yes No or Undecided
Male 0.452 0.103 0.555
Female 0.353/0.455 0.092/0.455= 1
Female 0.353 0.092 0.455 =0.7758 0.2242
Total 0.805 0.195 1 0.7758, y= 0
P (Y/X=1) 0,2242,
= y= 1
0 otherwise
• Biliefeness in Endemic for male
Belief in Endemic (Y)
0.8144, y=0
Gender Total
P (Y/X=0) = 0,1856, y= 1
(X)
Yes No or Undecided 0 otherwise
Male 0.452/0.555 0.103/0.555= 1 X=0 for male Y=0 for yes
=0.8144 0.1856 =1 for female =1 for No or undecided