REX
EDUCATION
STATISTICAL ANALYSIS
WITH SOFTWARE APPLICATIONS
John Paolo R. Rivera, Ph.D.
19.5601
17.9245
16.3672
14.0501
14.0162
12.5840
12.1480
10.7245
OUTCOMES-
BIZ-ACСТ BASED
Series EDUCATION First Edition
CONTENTS
Chapter 1. Definition of Statistics .3
Division of Statistics .3
1
2.
a. Descriptive Statistics .3
b. Inferential Statistics 4
FUNDAMENTALS 4
OF STATISTICS
3. Importance of Statistics.
4. Population and Sample. 4
5. Parameter and Statistic. .5
6. Quality of Data .6
7. Variables .6
a. Construct and Variable 7
b. Quantitative Variable and Qualitative Variable 7
c. Discrete Variable and Continuous Variable .8
d. Dummy Variable and Latent Variable 8
8. Levels of Measurement. 10
a. Nominal .11
b. Ordinal .11
c. Interval 12
d. Ratio 12
9. Classification of Statistical Procedures. 13
a. Parametric Statistics. .13
b. Nonparametric Statistics. .13
10. Types of Data .14
a. Cross-Section. .14
b. Time Series .14
c. Panel Data .17
11. Ethical Use of Statistics .19
xiii
1. Frequency Distribution .31
Chapter
Constructing Frequency Distribution. .31
2
DESCRIPTIVE
2.
3.
Cross Tabulations
Cumulative Frequency Distribution
.33
.36
a. Relative Frequency Distribution .36
STATISTICS
b. Absolute and Relative Cumulative
Frequency Distribution. .36
.37
4. Graphing Frequency Distribution
a. Histogram. .37
b. Charts and Graphs. .39
i. Bar Chart. .39
ii. Pareto Chart .40
iii. Pie Chart .41
iv. Time Series Graph .41
5. Misleading Charts and Graphs .42
6. Measures of Central Tendency (or Location) .43
a. Mean. .43
b. Arithmetic Mean, Weighted Mean,
and Geometric Mean .44
c. Median .47
d. Mode. .48
e. Relative Positions of Mean, Median, and Mode .49
7. Measures of Dispersion (or Variation). .49
.50
a. Range.
b. Mean Deviation (or Mean Dispersion) .51
c. Variance and Standard Deviation .51
d. Coefficient of Variation .54
8. Uses of Standard Deviation .55
a. Chebyshev's Theorem .55
b. Empirical Rule .56
9. Measures of Position .57
a. Quartiles .57
b. Deciles .58
c. Percentiles .58
xiv
10. Moments of Statistical Distribution .60
a. Mean and Variance. .60
b. Skewness .60
c. Kurtosis. .62
Chapter 1. Probability, Experiment, Outcome, and Event. .76
2. Conceptual Approaches to Probability .77
3
PROBABILITY
a,
b.
Objective Probability.
Subjective Probability
.77
.78
THEORY 3. Addition and Multiplication Rules of Probability .79
a. Addition Rule .79
b. Multiplication Rule 80
4. Contingency Tables. .82
Counting Rules .84
5.
a. Product Rule. .84
b. Permutation .84
c. Combination .85
d. Application of Counting Rules .86
6. Probability Distributions .87
Random Variables. .87
7. Discrete Probability Distribution .89
a. Mean, Varíance, and Standard Deviation of a
Discrete Probability Distribution .89
b. Binomial Probability Distribution. .90
c. Poisson Probability Distribution .92
8. Continuous Probability Distribution. .94
a. Mean, Variance, and Standard Deviation of a
Continuous Probability Distribution. .94
b. Uniform Probability Distribution 94
c. Normal Probability Distribution. .96
d. Standard Normal Probability Distribution .97
e. Exponential Probability Distribution .102
XV
Sampling Methods 118
Chapter 1.
Non-probability Sampling 119
4
a.
i. Purposive Sampling .119
il. Convenionce Sampling .119
SAMPLING .119
ii. Quota Sampling
DISTRIBUTION OF
THE MEAN b. Probability Sampling. .120
i. Simple Random Sampling .120
ii. Systematic Random Sampling. .121
iii. Stratified Random Sampling .121
iv. Cluster Random Sampling. .121
v. Multistage Random Sampling .121
2. Sample Size Determination 121
3. Sampling Distribution of the Sample Means. .125
4. Central Limit Theorem. .127
Chapter 1. Point Estimates and Intervai Estimates .142
5
a. Point Estimate .142
b. Interval Estimate and Confidence Interval .143
i. Confidence Interval
INFERENTIAL Estimation for Large Samples 144
STATISTICS
ii. Confidence Interval
Estimation for Small Samples .147
Iii. Confidence Interval
Estimation for Proportion .151
Iv. Finite Population Correction Factor .153
2. Hypothesis Testing 154
3. Statistical Hypothesis Testing Procedure .155
a. Null Hypothesis and Alternative Hypothesis. .156
b. One-Tailed Test and Two-Tailed Test .157
c. Level of Significance. .157
d. Type I Error and Type II Error. .158
e.. Test Statistic .158
f. Decision Rule: Critical Value Approach .158
g. Decision Rule: p-value Approach .159
xvi
Approaches to Hypothesis Testing .159
4.
a. Hypothesis Test about a
Population Mean (One-Sample, z Test) .159
b. Hypothesis Test about a
Population Mean (One-Sample, t Test) .163
c. Hypothesis Test about a
Proportion (One-Sample, z Test) .168
d. Hypothesis Test about a
Population Mean
(Independent Two-Sample, z Test). .169
e. Hypothesis Test about a
Population Mean (Independent Two-Sample, t Test)...172
f. Hypothesis Test about a
Proportion (Two-Sample, t Test) .174
5. Software Application in Inferential Statistics .176
a. Hypothesis Test for a Population Mean
(Independent Two-Sample,
Equal and Unequal Variances, t Test) .177
b. Hypothesis Test for a Population Mean
(Dependent Two-Sample, Paired Sample, t Test).....179
Chapter 1. The F Distribution and Analysis
6
of Variance (ANOVA) .194
a. Comparing Two Population Variances (F Test) .195
b. Analysis of Variance (ANOVA) Test .199
ANALYSIS OF c. Inferences about Pairs of Treatment Means. .203
VARIANCE
(ANOVA) AND 2. The Chi-Square Distribution and
.205
NONPARAMETRIC Nonparametric Hypothesis Testing
TESTS .206
a. Characteristics of the Chi-Square Distribution
b. Chi-Square Test Statistic and Analysis .206
c. Chi-Square Goodness-of-Fit Test:
Equal Expected Frequencies .207
d. Chi-Square Goodness-of-Fit Test:
Unequal Expected Frequencies .209
e. Chi-Square Test of Independence:
Contingency Table Analysis .211
f. Kruskal-Wallis H Test .214
g. Kendall WTest .217
xvii
h. Mann-Whitney U Test .220
i. Wilcoxon Signed-Rank Test (for Related Samples)... .226
3. Test for Normality .229
a. Jarque-Bera Test for Normallty. .230
b. Kolmogorov-Smirnov Test for Normality. .232
Chapter 1. Correlation Analysis .251
.253
Dependent and Independent Variables.
7 2. The Correlation Coefficient.
a.
b.
Pearson's Correlation Coefficient
Spearman's Rank Correlation Coefficient
.255
.255
.257
CORRELATION 3. Correlation and Causation .260
AND REGRESSION
ANALYSIS 4. Regression Analysis .262
Least Squares Principle .262
5. Simple Linear Regression. .263
a. Assumptions of Simple Linear Regression. .264
b. Estimating the Coefficients of
Simple Linear Regression Model .265
c. Estimating the Standard Error .267
d. Constructing the Confidence and
Prediction Interyals. .268
e. Coefficient of Determination. .270
f. Relationship of Correlation Coefficient,
Standard Error of Estimate, and
Coefficient of Determination. .271
6. Multiple Linear Regression .272
a. Assumptions of Multiple Linear Regression .273
b. Estimating the Coefficients of the
Multiple Linear Regression Model .274
c. Interpreting the Multiple Linear Regression .275
7. Regression Analysis with Dummy Variables.......280
8. Positive Monotonic Transformation of Data .282
9. Modeling Relationships of Multiple
Variables with Linear Regression .284
10. Advanced Regression Models .286
xviil
Appendices .297
1. List of Statistical Tables. .297
a. Cumulative Normal Distribution .297
i. Table of Standard Normal Probabilities
for Negative Z-scores .297
ii. Table of Standard Normal Probabilities
for Positive Z-scores .297
b. Critical Values of the t Distribution .297
C. Critical Values of the FDistribution. .297
i. Critical Values of F at the 5% Significance Level .. .297
ii. Critical Values of Fat the 1% Significance Level ...297
iii. Critical Values of F at the 0.1%
Significance Level. .297
d. Critical Values of the Chi-Square Distribution. .297
e. Critical Values of the Mann-Whitney U. .297
f. Critical Values of the Wilcoxon Signed-Rank Test.....297
g. Critical Values of the One Sample Kolmogorov-Smirnov
Test of Normality. .297
h. Critical Values for the Spearman's Rank
Correlation Coefficient .297
2. Appendix 1 .298
3. Appendix 2 .299
4. Appendix 3 .300
5. Appendix 4 .306
6. Appendix 5 .307
7. Appendix 6 .308
8. Appendix 7 .309
9. Appendix 8 .310
10. Appendix 9: Rubric for Grading Concept Map
and Essay Questions 311
11. Sample Datasets. .312
a. World/International. .312
b. Philippines/Local .312
Index .314
Worksheets .322
xix
FUNDAMENTALS
CHAPTER
11
OF STATISTICS
Chapter Goals and Learning Outcomes
After completing this chapter, you are expected to:
list the importance of studying statistics;
distinguish between parameter and statistics;
differentiate between descriptive and inferential statistics;
use the various types and kinds of variables and data in an experiment;
differentiate the four levels of measurement of data; and
state the conditions when to use a specific statistical procedure from the kind of data
on hand and its corresponding level of measurement.
INTRODUCTION
More advanced technology has enabled never-before-seen data collection methods
and given us access to a far larger volume of data. Our ability to gather and compile
information has become more efficient. The heads of our national government institutions
(e.g., Philippine Statistics Authority, Bangko Sentral ng Pilipinas, Philippine Institute for
Development Studies, National Economic and Development Authority, and Philippine Stock
Exchange, among others), as well as international and multilateral organizations (e.g., Asian
Development Bank, ASEAN Secretariat, World Bank, World Trade Organization, and World
Travel & Tourism Council, among others), are proficient in handling and analyzing data.
They are cognizant of the critical importance of statistical techniques that have the capability
to affect markets, stimulate the economy, and provide direction to policy making.
The importance of information and knowledge produced by cutting-edge statistical
processes can no longer be overstated. As Peter Drucker, a management consultant,
educator, and author, whose works contributed to the philosophical and practical foundations
of modern management, is often quoted for, "If you cannot measure it, you cannot improve
it" (MacKenzie, n.d.). We can construe that what you cannot measure, you cannot manage
and grow (Zak 2013). Thus, the statement of economist and former chairman of the US
Federal Reserve, Alan Greenspan, makes sense in today's environment given globalization,
the expansion of data availability, and the growing demand for evidence-backed judgments.
To be competitive, "Workers must be equipped not simply with technical know-how, but also
with the ability to create, analyze and transform information, and to interact effectively with
others" (The Federal Reserve Board 2000).
With the growing misinformation and disinformation (i.e., the spread of "fake news"),
it is our responsibility to separate facts from opinions and reality from fiction (i.e., people's
figments of imagination). We must arrange the facts correctly, carefully examine the data
and information, and spread accurate and impartial knowledge.
Thus, in this chapter, we look into the world of statistics and its jargon, its types and
emphasis on data, measurement, management, and ethical aspects. Being familiar with the
fundamental concepts of statistics as a discipline and data management would be helpful
as you advance to the succeeding chapters. All of these will be repeatedly mentioned and
applied as we go through each lesson. It is recommended that before you end this chapter,
work on the individual, group, data, and linking concept exercises. Also, perform self-
assessment and diagnose through the checklist whether you have sufficiently prepared for
more adventures in the next chapters:
Chapter 1 | Fundamentals of Statistics 1
Diagnostic Exercises
General Instructions: Answer the following questions to prepare you for this chapter. This
exercise will assess your prior knowledge and familiarity with the concepts to be covered in
this chapter. As much as possible, answer this on your own. You may check your answers
afterward with your teacher and classmates.
1. Short-Answer Items:
a. The area of statistics concerned with the organization, summarization, and
presentation of data in an informative manner is called
b. The area of statistics that validates a statement about a population using sample
data is called
c. The metric that represents an aspect of the population of interest is called
d. The metric that represents an aspect of a sample is called
e. A variable that is measured using whole numbers only with definite gaps between
values is called
f. A variable that is measured by any rational number is called
g. When specific assumptions about the distribution of the underlying population where
the sample was sourced are required, the appropriate statistical approach to use is
h. When no assumptions about an underlying population distribution is required, the
appropriate statistical approach to use is
i. Whena data representing a variable is categorical in nature, its level of measurement
is
j. If the data represents a ranking, its level of measurement is
2. Binary-Choice Items: If the statement is correct, write TRUE. If the statement is incorrect,
write FALSE and explain why by stating what made it false and what will make it true.
a. Interval data has a true zero value.
b. The college where you belong in your university is a latent variable.
c. Cross-section data is measured for different individuals across different time periods.
d. A construct is what we use to measure a variable.
e. A sample is drawn from a proportion of the population.
2 Statistical Analysis with Software Applications
DEFINITION OF STATISTICS
A powerful tool to create, understand, and digest information is statistics. "Statistics
is the science of collecting, organizing, summarizing, presenting, analyzing, and
interpreting numerical information from data to assist in making more effective
decisions" (Brase and Brase 2010, 4; Lind et al. 2006, 5) in the face of volatilities,
uncertainties, complexities, and ambiguities. We use statistics in our daily lives, with or
without us knowing it. For instance, we are often faced with decisions like whether to
bring an umbrella or raincoat or none; to drive public transportation; to choose
or to take
numbers to bet on in the lottery; and to choose which vaccine is best for us. Moreover, in
grocery stores and, malls, we like taking samples (i.e., "free tastes") of products first before
purchasing them; we test light bulbs, appliances, and equipment before paying for them;
among others. Meanwhile, companies regularly take samples of their products to test for
quality. On a national scale, political candidates research the distribution of their supporters
in different regions and bailiwicks; the government, among other things, selects the brand
of vaccination to be given depending on efficacy. These situations call for decision-making
processes that must be anchored not only on the decision-maker's values and attitudes but
also on available information. Statistics and its methods help us examine the information on
hand. This is the very reason why you are taking a statistics course.
DIVISION OF STATISTICS
Statistics can be categorized into two: descriptive and inferential statistics.
Descriptive Statistics
From our definition of statistics earlier, the area of statistics concerned with the
organization, summarization, and presentation of data in an informative manner is called
descriptive statistics.
Examples of descriptive statistics include but are not limited to:
reports based on the average of all weather stations in the Philippines, excluding
Baguio, that the mean annual temperature is 26.6 degrees Celsius and that
the coolest month falls in January with a mean temperature of 25.5 degrees
Celsius, while the warmest month occurs in May with a mean temperature of 28.3
degrees Celsius (Philippine Atmospheric Geophysical and Astronomical Services
Administration, n.d.);
unemployment in the Philippines in June 2022 is at 6 percent, equivalent to 2.99
million Filipinos (Mercado 2022); and
the Philippine auto industry sold 142,742 units of motor vehicles during the first
half of 2021 (Laurel 2021).
Chapter 1 | Fundamentals of Statistics 3
Inferentlal Statistice
Meanwhile, by going beyond descriptive to claim something aboutapopulation from
a sample taken trom that population is called Inferential statistlcs, which is also known as
inductive statistics or statietical inforence.
Examples of statements using Inferentlal statistics include but are not limited to:
findings saying that the Philippines ranked lowest in mathematics and science
tests for Grade 4 students, among 58 countrles, using a sample of Grade 4 Filipino
students (Mendoza 2020); and
reports saying the Philippines suffered a seven-point drop in the English Proficiency
Index, using a samplo of 2.2 millon adults from 100 countries and territories in
2019 (CNN Philippines Staff 2020).
IMPORTANCE OF STATISTICS
From ourdefinition of statistics, we can construe that it is the science of learning from
data. It will teach you the appropriąte methods to collect data, use the correct data analytic
tools, and effectively present the results to potential users. Therefore, statistics is vital in
creating new knowledge gleaned from data, making effective decisions that are backed
by data, and making roasonable proedictions based on the behavior of data. As a result,
you understand a phenomenon, or a subject matter much better with a higher degree of
objectiveness.
In today's society, learning statistics is crucial. First, learning from data allows you to
navigate common personal and professlonal problems thatcan lead you to wrong conclusions
if you only use subjective judgment (1.e.,your emotions). Second, the abundance of choices
available to us in modern life forces us to critically assess our options in order to choose the
course of action that will best serve our needs. The lessons from statistics will help us have
the rigor in making effective decisions.
POPULATION AND SAMPLE
From our definition of inferential statistics above, we also need to highlight the
definitions of population and sample. On one hand, the population (denoted by M) is the
entire set or totality of individuals or objects of interest. We will formally define an individual
later. On the other hand, the sample (denoted by n) is a part or subset of the population of
interest. Figure 1.1 illustrates the relationship between population and sample.
The reason sample instead of the whole population is economic. Because
for taking a
the cost of conducting a census is exorbitant and prohibitive, taking a representative sample
is practical and feasible. That is why a census is done only at certain periods while other
4 Statistical Analysis with Software Applications
surveys are done regularly. This is the same reason why chefs do not eat the entire dish
they prepare to know if it tastes good or not; why you do not drink the entire bottle of red
wine to know if it suits your palate; why you do not eat the entire cauldron of soup to know if
it is delicious or not; and why a nurse or a medical technician will not extract all your blood
to conduct your blood tests. Getting a sample from the entire thing (i.e., population) makes
sense.
Population
1 2 3 4 5 6 7 8 9 10 11 12
2 5 8 11
Sample (every 3rd)
Figure 1.1. Population vs, Sample
PARAMETER AND STATISTIC
One must also know whether the data on hand is composed of all individuals of interest
(population data) selected individuals only (sample data). Knowing this will allow
or some
you to determine whether you are working with a parameter or a statistic.
On one hand, an important aspect of the population is represented by a parameter,
which is a metric. On the other hand, a statistic (notice that the word is singular) is a metric
that depicts a feature of a sample taken from the target population.
For example, consider the tourists who visited Siargao Island. Assuming we have data
from all the tourists who visited the destination in the past year, then we have population
data. The average age of all tourists who visited the destination last year is a parameter.
Similarly, the proportion of females, in the population of all tourists who visited the destination
last year is also a parameter.
However, if our data is sourced from some of the tourists, we have sample data.
Correspondingly, the average age of tourists who visited the destination last year is
represented by the sample, which we call a statistic. Likewise, the proportion of female
tourists in the sample is a statistic.
It important to know that if a differeht sample is used, it may yield different values for
is
the average age and proportion of female tourists. Hence, a population parameter is fixed
for a given population, while a sample statistic can vary for different samples.
Chapter 1 | Fundamentals of Statistics 5
QUALITY OF DATА
Your intuition, built-in system of inference, and bounded rationality will all be
strengthened by the statistical techniques we will explore. In other words, the results of your
chosen statistical methods should confirm, explain, or guide our sound common sense. While
we must realize that statistical methods cannot create miracles, they can assist us inmaking
decisions, but not all conceivable decisions. As early as now, I must stress that a properly
used statistical technique is only as accurate as the information, facts, or data used (i.e.,
quality in, quality out or QIQO; garbage in, garbage out, abbreviated as GIGO). Likewise,
statistical results must be read and interpreted by someone who knows the methods and
the context in which they have been applied. This is similar to medical laboratory results
indicating that the "results are best interpreted by a healthcare professional and are not
intended to be used as the sole means for diagnosis or management."
Data are critical elements particularly in scientific research. They are the bases
for understanding and analyzing the questions, inquiries, and problems, which could
subsequently provide solutions in a discipline, an economy, a society, or a community. Thus,
the value of quality data must be emphasized. When quality data is'used, the information
generated by data analysis becomes more meaningful and valuable to the users.
Data collection is anchored on the principle of "due professional care." It must be hinged
on the accumulation of adequate empirical evidence to render claims about an issue or
subject matter supported. Of equal importance, data must be read, analyzed, and interpreted
so that meanings can be formulated, interpolated, and extrapolated. It is dangerous to make
personal, professional, or business decisions based on poor data as it may be disastrous.
Hence, it is also important to be cognizant of your data source and its nature, as well as
to address the data's credibility problems
(Redman 2013). You have to ensure the credibility
of the people or organizations that gathered the data. It is also essential to ask how and
when the data was gathered. It is important to know the existence of credibility issues to
compel users to analyze data with care, mitigațe hasty generalization, and address areas of
concern to improve data quality. This is also a call for everyone to focus on improving the
way new data is created rather than exerting massive effort to clean up existing bad data.
VARIABLES
The fundamental requirement to conduct statistical analysis and to make decisions is
data gathering. To do this, there is a need to identify the individuals to be involved in a study
and their characteristics that are of interest.
Individual is the unit of analysis in a certain study, which can be in the form of people.
places, objects. The characteristic of the individual being measured is called a variable.
or
For example, if we want to conduct a study about the tourists who visited Siargao Island
last year, then the individuals in the study are those who visited the destination last year for
6 Statistical Analysís with Software Applications
leisure purposes. One variable might individuals. Other variables can be
be the age of the
nationality, sex, gender, monthly income of tourists, their travel budget, and their number of
days of stay, among others. Regardless of the variable we use, it would be meaningless to
include information from people who have not visited the destination.
Construct and Variable
In assigning a variable, we need to first understand a construct. A construct is a
concept, idea, abstract thing, or attribute that exists in our human brain, is not directly
observable, is founded on specific ideas, and theoretically represents real objects and
processes. It can be represented or measured by a variety of variables. Consaquently, a
variable can also be described as the operational definition or measurable representation of
a construct. Correspondingly, a construct is the conceptual definition of a variable.
For example, a person's intelligence is a construct. It may be measured by any of
the following variables: a person's intelligence quotient (IQ) score, emotional quotient (EQ)
score, grade point (GPA), cumulative GPA (CGPA), or score on a scholastic aptitude
average
test (SAT), among others. Another example of a construct is beauty. While some famous
adages say that "beauty is subjective" and "beauty is in the eye of the beholder," people
have attempted to measure or quantitatively represent it using the golden ratio, also known
as the Vitruvian ratio or the ratio of beauty expressed as the mathematical ratio of 1:1.618;
rating scales indicating the degree of concurrence about an individual being beautiful; and
the number of wins in pageants, among others. Popularity is also example of a construct,
an
which can be represented by the number of likers, subscribers, and supporters. The list of
constructs and corresponding variables is endless. You need to justify why you think your
chosen variable is appropriate to represent your construct.
Quantitative Variable and Qualitative Variable
Notice from our examples above that the possible variables we have listed can be
qualitative or quantitative in nature. They are the two basic typeş of variables.
If a variable is reported numerically, then it is a quantitative variable. Examples are:
the number of members in your family, the number of pets you have at home, the mileage
of your car, the number of COVID-19 vaccine doses administered in the Philippines, the
population of Metro Manila, the net income of a fast food chain, the balance of your savings
Island,
account, the value of stocks in your portfolio, the number of tourists visiting Boracay
and the foreign exchange rate between the Philippine Peso (PHP) and the US Dollar (USD),
among others.
Meanwhile, a qualitative variable is also called an attribute, which is a non-numeric
characteristic of individuals. Examples of qualitative variables are gender, nationality, highest
educational attainment, marital status, race, religious affiliation, employment status, type of
house owned, type of automobile owned, place of birth, income class, and political affiliation,
among others. Notice that they are mostly demographic variables or attributes that are
studied statistically and whose interest is primarily in how many are under each category.
Chapter 1 | Fundamentals of Statistics 7
Discrete Variable and Continuous Variable
Quantitative variables can also be categorized either as continuous.
discrete or
Discrete variables are expressed, measured, or represented by whole numbers only with
definite gaps between values. Examples are:
the number of cars in a dealer's showroom;
bedrooms in a condominium unit (e.g., one bedroom, two bedrooms, or three
bedrooms);
the number of motor vehicles using the South Luzon Expressway (SLEX) daily
(e.g., 300,000 to 500,000 motor vehicles); and
the number of signatory countries in the Paris Climate Agreement, World Trade
Organization, or Regional Comprehensive Economic Partnership, among others.
On the other hand, continuous variables can have any value within a specific range.
Examples are:
your weight and height;
your house's electricity bill;
weight of your package for shipping from Manila to Davao;
duration of flights from Clark to Caticlan;
distance from Cebu to Seoul;
price of laptops sold;
inflation rate of the Philippines;
cost of roll-on/roll-off (RORO) tickets from Boac, Marinduque to Lucena, Quezon;
total contract price of your lot in Laguna; and
outstanding balance of your credit card bill, among others.
In short, discrete variables result from counting while continuous variables result from
measuring.
Dummy Variable and Latent Variable
In order to process and make sense of qualitative variables, we can construct a dummy
variable, also known as an indicator variable, design variable, contrast, one-hot coding, and
binary basis variable, which is a numeric variable that represents categorical data, such as
those mentioned under qualitative variables.
Dummy variables are dichotomous quantitative variables. They can take on only two
quantitative values, say, either 1 or 0. Simply, values indicate the presence or absence of
something. Usually, 1 represents "yes" and 0 otherwise.
For example, sex as a variable has two categories: either male or female. Suppose
that we assign female category. There is no universally accepted
as our base or reference
approach to choose the reference category (Grace-Martin, n.d.). In practice, the base
category or reference category is the one that will be dropped. In coding, plotting, and
analyzing dummy variables,,we follow the rule D-1, which means we deduct 1 from the
8 Statistical Analysis with Software Applications
total number of categories. Because sex categories-male and fernale, the base
has two
category is hidden from sight or dropped because it is captured when the value assumed is
0. It is done to avoid the dummy variable trap (Gujarati et al. 2017).
Therefore, if an individual is male, he will take a value of 1, and 0 otherwise, if she is
female. On the other hand, nobody is stopping us from assigning male as the base category.
Therefore, if an individual is female, she will take a value of 1, and 0 otherwise, if he is male.
See Table 1.1 for the illustration. Shaded column is the dropped category.
Table 1.1. Illustrating Coding for Dummy Variables (Sex)
Individual Male Female (as base category)
Andres Bonifacio 1 이
Gabriela Silang 이 1
Antonio Luna 1 이
Macario Sakay 1 0
Gregorio Del Pilar 1 이
Jose Rizal 1 이
Melchora Aquino 이 1
Josefa Llanes Escoda 0 1
Apolinario Mabini 1 0
Maria Orosa 이 1
Individual Male (as base category) Female
Andres Bonifacio 1 0
Gabriela Silang 이 1
Antonio Luna 1 0
Macario Sakay 1 0
1 이
Gregorio Del Pilar
Jose Rizal 1 0
Melchora Aquino 이 1
Josefa Llanes Escoda 이 1
Apolinario Mabini 1 이
Maria Orosa 이 1
For dummy variables with more than two categories, like island groupings of the
Philippines (i.e., Luzon, Visayas, and Mindanao), the same coding instructions apply. As an
example, we code the location of the home province of Philippine presidents according to
island groupings. See Table 1.2 for the illustration with Mindanao as the base category. The
shaded column indicates the dropped category. It follows when Luzon or Visayas is set as
the base category.
Chapter 1 | Fundamentals of Statistics 9
Table 1.2. Illustrating Coding for Dummy Variables (Island Grouping)
Philippine President Home Province
Mindanao (as
tas per common (as per common Luzon Visayas
base category)
knowledge) knowledge)
Emilio F Aguinaldo Cavite 1 이 이
Manuel L Quezon Aurora 1 이 이
José P Laurel Batangas 1 이 이
Sergio Osmeña, Sr. Cebu 이 1 이
Manuel A. Roxas Capiz 이 1 0
Elpidio R. Quirino Ilocos Sur 1 이 이
Ramon F. Magsaysay Zambales 1 이 0
Carlos P. Garcia Bohol 이 1 이
이
Diosdado P. Macapagal Pampanga 1 0
Ferdinand E. Marcos, Sr. llocos Norte 이 이
Corazon C. Aquino Tarlac 1 이 0
1 이 이
Fidel V. Ramos Pangasinan
Metro Manila 1 이 이
Joseph Ejercito Estrada
Gloria Macapagal-Arroyo Pampanga 1 이 이
Benigno S. Aquino III Tarlac 1 이
0
Rodrigo Roa Duterte Davao City 0 1
Ferdinand R. Marcos, Jr. Ilocos Norte 1 0 이
In statistics, we also have latent variables, also known as hidden variables, which
are indirectly observed or inferred using other variables that can be directly observed.
They represent abstract concepts like behavioral states, mental states, or data structures.
Examples are extraversion, conservatism, quality of life, happiness,
spatial ability, wisdom,
morale, and business confidence, among others. All of them cannot be directly measured.
However, by linking them to other observable variables, their values can be inferred from the
measurements of the observable variables. Take the case of quality of life. Since it cannot
be measured directly, observable variables, such as employment, wealth, income, leisure
time, education, environment, physical health, mental condition, and social belongingness,
among others, are used to infer the value of quality of life.
Detailed discussions of latent variables and how to treat them are found in more
advanced statistical textbooks.
LEVELS OF MEASUREMENT
Understanding the corresponding levels of measurement is essential now that
variables can be categorized as quantitative, qualitative, discrete, or continuous. In studying
statistics, it is integral because knowing a variable's level of measurement will guide us on
the appropriate treatment that must be done to the data, which I willdiscuss in detail later.
The four levels of measurement of a variable are: nominal, ordinal, interval, and ratio.
10 Statistical Analysis with Software Applications
Nominal
Nominal is applied to data that consists of categories. There is no order or sequence
tothe categorization, nor is there a prescribed criteria by which data can be sequenced.
That is, they are observations of a qualitative variable that can be counted and classified.
Examples include those we have enumerated under qualitative and dummy variables.
Note that categories are mutually exclusive and exhaustive. They are mutually
exclusive when an observation for an individual can only be classified into a single category.
They are exhaustive when an observation for an individual must be classified in one of the
categories.
Ordinal
Ordinal is applied to the data that can be ranked or sequenced ín order. With ordinal
data, we cannot distinguish the magnitude of the difference between the values of data. That
is, there is no meaning in the differences between the values of data. Examples include the
ranking of the students class; the hierarchy of positions in a company (e.g., entry level,
in a
associate, coordinator, supervisor, manager, director, and executive/management); the alert
level status of the Philippines with respect to coronavirus (COVID-19) cases (e.g., Alert
Levels 1, 2, 3, 4, and 5; see Table 1.3) (Gregorio 2021); PAGASA's public storm warning
signals (e.g., Signal Nos. 1, 2, 3, 4, and 5; see Table 1.4) (PAGASA, n.d.), and country
standings in terms of medals earned in the Tokyo Olympics 2020, among others.
Table 1.3. COVID-19 Alert Levels as an Example of Ordinal Data
Alert Level Indicator/Triggers/Remarks
Virus transmission (cases) is low and decreasing; total bed utilization rate and ICU utilization
1 rate are low; 70 percent of senior citizens, people with comorbidities, and eligible population
have been vaccinated.
Virus transmission (cases) is low and decreasing; healthcare utilization is low, or cases are
2 low but increasing, or cases are low and decreasing, but bed utilization and ICU utilization is
increasing.
Virus transmission (cases) is high and/or increasing and there is increasing utilization of
3
hospital beds and ICUs.
4
Virus transmission (cases) is high and/or increasing, and hospital bed and ICU utilization is
high.
5 Virus transmission (cases) is "alarming" and hospital bed and ICU utilization is at critical levels.
Source: Culled from Gregorio (2021).
Table 1.4. Public Storm Warning Signals as an Example of Ordinal Data
Lead Time (in
Public Storm Wind Speed
hours) on First Impact of the Wind
Warning Signal (in kph)
Issuance Only
1 36 30-60 No damage to very light damage
2 24 61-120 Light to moderate damage
3 18 121-170 Moderate to heavy damage
4 12 171-220 Heavy to very heavy damage
5 12 >220 Very heavy to widespread damage
Source: Philippine Atmospheric Geophysical and Astronomical Services Administration.
Chapter 1 | Fundamentals of Statistics 11
Interval
Interval is applied to data that can be ranked or sequenced (i.e., includes the
data), but differences between data values are constant and
characteristics of ordinal
have meaning. Examples include: temperature, shoe size, IQ scores, and years in which a
congressman was elected to the House of Representatives.
understand it better, take temperature as an example. Suppose the high
To
temperatures on three successive summer days in Angeles, Pampanga are 38, 41, and 35
degrees Celsius. These values can easily be ranked, but we can also differentiate among
the given temperature values because 1 degree Celsius is a constant measurement unit.
There are also equal differences between and among temperature values regardless of their
location on simply, the difference between 41 and 38 degrees Celsius is
the scale. To put it
3, and the difference between 38 and 35 degrees Celsius is also 3. It is also vital to mention
that 0 is just a point on the scale, which does not indicate the absence of the condition.
With temperature, 0 degrees does not mean the absence of temperature; it means that it is
freezing cold.
Consider another example to better understand interval level. Take the years in which
a congressman was elected to the House of Representatives, say 2016, 2019, and 2022,
as an example. It is interval because the years can be ordered, and the difference between
the years has meaning. For instance, 2022 ís six years later than 2016. Similarly, 2022 is not
twice as large as year 1011. Likewise, year 0 does not mean the absence of a year.
Ratio
Ratio is applied to data that [Link] ordered, and differences between data values and
ratios of data values have meaning (i.e., includes the characteristics of interval data), and
data have a true zero (i.e., the existence of a true zero distinguishes ratio from interval). Age,
height, weight, Kalibo and Caticlan distance, salaries and wages, money in your savings
account, number of cases, active cases, deaths, recoveries from COVID-19, quantity of
bags sold at the department store, and your taxable income for 2022 are a few examples.
The characteristics of the four levels of measurement are summarized in Table 1.5.
Table 1.5. Characteristics of the Four Levels of Measurement
Characteristics Levels of Measurement
Nominal Ordinal Interval Ratio
Data classifications are mutually exclusive and exhaustive. ✓ ✓
Data classifications are ordered according to the amount of
X ✓ ✓
the characteristic they possess.
Equal differences in the characteristic are represented by
equal differences in the measurements.
The zero point Is the absence of the characteristic. x X X
Source: Culled from Lind et al. (2006)
While identifying the level of measurement of your data may sometimes be confusing
you may also use Table 1.6 to guide you. It takes patience and practice to seamlessly
identify the data's level of measurement.
12 Statistical Analysis with Software Applications