MODULE 1
EDUCATIONAL STATISTICS
LESSON 1: INTRODUCTION TO STATISTICS
INTRODUCTION
Overview of Introduction to Statistics
Introduction to Statistics provides the foundational concepts and tools for
collecting, organizing, analyzing, and interpreting data. It equips learners with methods
for summarizing data through descriptive measures, making inferences from sample
data, and applying probability theories to decision-making. This introductory knowledge
serves as the stepping stone for more complex statistical techniques and fosters
analytical thinking essential in academic and professional research (Triola, 2020; Uyanik
& Güler, 2020).
Relevance of the Topic in Advanced Statistics
The principles learned in an introductory statistics course are crucial for understanding
advanced statistical concepts such as multivariate analysis, regression modeling, time-
series forecasting, and machine learning algorithms. Without mastery of basic statistical
reasoning, students may struggle to interpret advanced models or assess the validity of
statistical findings. Advanced Statistics relies heavily on skills such as hypothesis testing,
confidence interval construction, and understanding sampling distributions—all rooted in
introductory concepts (Ben-Zvi & Garfield, 2020; Bluman, 2021).
Connection to Prior Knowledge
Learners entering statistics courses often have prior exposure to mathematics
topics like algebra, proportions, and basic data representation (e.g., bar graphs, line
charts). This prior knowledge facilitates comprehension of statistical measures such as
mean, median, standard deviation, and correlation. Furthermore, skills in logical
reasoning, problem-solving, and critical thinking—often developed in earlier mathematics
or science subjects—help students grasp statistical arguments and apply them in real-
world contexts (Franklin et al., 2020; Ridgway, 2021).
Gender Perspective: The Role of Gender in Statistics
From a gender perspective, statistics helps identify and address disparities in areas like
education, health, employment, and politics. Gender-disaggregated data enables
policymakers to monitor equality and design evidence-based interventions, such as
addressing wage gaps and women’s underrepresentation in certain industries.
Incorporating gender analysis ensures data interpretation is inclusive and equitable
(Buvinic et al., 2020; United Nations, 2021).
By the end of this lesson, students shall be able to:
1
By the end of this lesson, students shall be able to:
1. write the definition of Statistics;
2. discuss the two branches of Statistics including the common tools used;
3. explain the research process;
4. state the goals of analysis;
5. identify variables and their functions;
6. explain & determine the levels of measurement;
7. describe the role of Statistics in research;
8. determine qualitative & quantitative variables;
9. identify whether the data is continuous or discrete, and;
10. express the appreciations & importance of statistics in the chosen
course/field.
How to understand the future of the
disease? How to know how many How long it will take to complete a
numbers of people are suffering from the project?
disease like COVID 19? How many have
died? Healed?
How do we determine the quality of a Have you ever seen weather
certain product that we are about to forecasting? Do you know how the
purchase? government does the weather
forecasting?
2
How we can predict any natural disaster How to determine the winners in an
that may happen shortly? How to help the election like president of the country?
rescue team to do the preparation to
rescue the life of the people who are in
danger?
How much premium to pay for your How to calculate which consumer
insurance? How the insurance companies goods are available in the store or not?
decide the premium amount?
How to find out which store needs the
consumer goods & when to ship the
products?
How to calculate all the stock prices in the How to help the sport persons to get the
financial market? idea about their performances in the
particular sports?
What is Statistics?
Statistics is an art and a branch of science which deals with the
collection, organization, presentation, analysis and interpretation
of data.
Recorded data such as the number of business permits issued,
number of customers eating at a restaurant, number of vaccines
formulated, number of functional hospitals, number of medicines, the
size of enrollment at SUNN, and so on.
3
Numerical characteristics calculated for a set of data (e.g., mean,
median, mode).
The backbone of RESEARCH.
Two Branches of Statistics
1. Descriptive Statistics
- is concerned with the collection, organization and presentation
of data in a form understandable to all.
- Its objective is to summarize some of the important features of a set
of data.
- Commonly used descriptive statistics:
Item Reason
Summarize categorical or numerical
a. Frequency & percentages
data counts.
Visual summaries (bar charts,
b. Graphic presentation of the data
histograms, pie charts, etc.).
Organizes data into classes with
c. Frequency distribution
corresponding counts.
Mean, median, mode — describe
d. Measures of central tendency
data’s center.
Range, variance, standard deviation
e. Measures of variability
— describe spread.
f. Measures of skewness and Describe shape and tails of the
kurtosis distribution.
g. Measures of position or relative Percentiles, quartiles, deciles —
standing position relative to dataset.
Z-scores and curve interpretation
h. Standard scores and the normal
are descriptive ways to standardize
curve
and compare data.
2. Inferential Statistics
- is concerned with the formulation of conclusions or
generalizations about a population based on an observation of a
sample drawn from a population.
- It is relevant when the researcher wishes to generalize findings
from a sample to a population or from a smaller group to a larger
group.
- it aims to give information about large groups of data without
dealing with each and every element of these groups.
- Inferential statistics:
a. Parametric (interval or ratio data is used & population is
normally distributed)
a.1. z-test
4
- test of difference bet. sample & population means
a.2. t-test
- test of difference in 2 groups
a.3. analysis of variance (ANOVA)
- test of difference in 3 or more groups
a.4. analysis of covariance
- combines ANOVA & regression analysis
a.5. multivariate analysis of variance
- with 2 or more dependent variables
a.6. pearson’s t
- test of correlation
b. Non parametric (nominal or ordinal data is used & population
is not normally distributed)
b.1. chi-square tests
- test of association bet. categorical variables
b.2. sign test
- test of 2 related samples
b.3. mann-whitney u test or wilcoxon rank-sum test
- test of 2 independent groups
b.4. kruskal-wallis h test
- test of 3 or more independent groups
b.5. friedman’s test
- test of 3 or more related groups
b.6. spearman rho
- test of correlation
The Research Process
Why do research?
Formulate the problem (SMART)
- Specific, measurable, attainable, relevant, & time bound.
Define the population of the study
Population – all subjects under investigation
- the set of all elements of interest in a particular
study
Sample – a subset of the population
Identify the variable/s of the study
Variable – measurable characteristics of the subject
-any entity that can take on different values
Example:
Problem:
5
What is the mean weekly allowance of a NONESCOST
Statistics student for the first semester of A.Y. 2018 – 2019?
Population of the study:
ALL NONESCOST Statistics students for the first semester, A.Y.
2018 – 2019.
Variable/s:
Weekly allowance of a Statistics student
(Anticipated) Conclusion:
The mean weekly allowance of a NONESCOST Statistics
student for the first semester of A.Y. 2018 – 2019 is
____________.
Goals of Analysis
- directed toward finding out from the data one or more of the following
attributes or characteristics of the group being studied:
1. Central tendency – general characteristic of the group
Examples:
a. To determine the mean weekly allowance of SUNN freshmen
for the first semester, A.Y. 2025 – 2026.
b. To determine the percentage of SUNN students who prefer a
Apple over a Samsung cellphone.
2. Variance in the group – how individual members of the group vary
from the average characteristics of the group.
Examples:
a. To determine the age range of the students in this Stat class.
b. To determine if the Stat final grades in this class are similar.
3. Difference within the group/between groups – whether or not
subgroups of the group/two separate groups being studied are
different or similar on certain traits investigated.
Examples:
a. To compare the mean no. of coke sakto bottles consumed in
a week between the male and female SUNN students.
b. To determine if there is a significant difference in the mean
number of text messages sent in a day among the students
from the eight different colleges of SUNN.
4. Relationships within the group – if relationship between certain
variables covered in the study exist.
Examples:
6
a. To establish if there is a significant relationship between
choice of cellphone brand and the college a SUNN student
belongs to.
b. To determine if relationship status and final grades in Stat
are independent.
5. Prediction – establishing a mathematical/statistical model to predict
future outcomes.
Examples:
a. What factors influence a graduate’s ability to land a job within
one year after graduation?
b. What is the estimated sale of a particular restaurant for next
week if the present conditions hold?
Types of Variables:
1. Categorical
- Attributes are in terms of categories.
- Example:
Sex: Male
Female
2. Numerical
- Attributes are in terms of counts or measurements.
- Distinctions:
a. Discrete Variable
-Uses the process of counting to generate data
-values of attributes are in terms of whole numbers only
-Examples:
a. Number of t-shirts owned
b. Number of pocketbooks read
b. Continuous Variable – use measuring instrument
-uses the process of measuring to generate data
-values of attributes may have fractional or decimal parts
-Examples:
a. weight of a package
b. volume of water
Functions of variables:
- Important if the investigation is about cause & effect.
- Distinctions:
a. Independent Variable
-what the researcher (or nature) manipulates – a
treatment or program or cause
b. Dependent Variable
7
-what is affected by the independent variable – the effects
or outcomes
- Example:
Study/Problem: the effects of a new educational program on
student achievement.
Independent variable – the program
Dependent variables – measures of achievement
Dependent and Independent Variables
An independent variable, sometimes called
an experimental or predictor variable, is a variable that is being manipulated in an
experiment in order to observe the effect on a dependent variable, sometimes called
an outcome variable.
Imagine that a tutor asks 100 students to complete a math test. The tutor
wants to know why some students perform better than others. Whilst the tutor does
not know the answer to this, the tutor thinks that it might be because of two reasons:
(1) some students spend more time revising for their test; and (2) some students are
naturally more intelligent than others. As such, the tutor decides to investigate the
effect of revision time and intelligence on the test performance of the 100 students.
The dependent and independent variables for the study are:
Dependent Variable: Test Mark (measured from 0 to 100)
Independent Variables: Revision time (measured in
hours) Intelligence (measured using IQ score)
The dependent variable is simply that, a variable that is dependent on an
independent variable(s). For example, in our case the test mark that a student
achieves is dependent on revision time and intelligence. Whilst revision time and
intelligence (the independent variables) may (or may not) cause a change in the test
mark (the dependent variable), the reverse is implausible; in other words, whilst the
number of hours a student spends revising and the higher a student's IQ score may
(or may not) change the test mark that a student achieves, a change in a student's
test mark has no bearing on whether a student revises more or is more intelligent
(this simply doesn't make sense).
Therefore, the aim of the tutor's investigation is to examine whether these
independent variables - revision time and IQ - result in a change in the dependent
8
variable, the students' test scores. However, it is also worth noting that whilst this is
the main aim of the experiment, the tutor may also be interested to know if the
independent variables - revision time and IQ - are also connected in some way.
In the section on experimental and non-experimental research that follows, we
find out a little more about the nature of independent and dependent variables.
Experimental and Non-Experimental Research
o Experimental research: In experimental research, the aim is to manipulate an
independent variable(s) and then examine the effect that this change has on a
dependent variable(s). Since it is possible to manipulate the independent variable(s),
experimental research has the advantage of enabling a researcher to identify a
cause and effect between variables. For example, take our example of 100 students
completing a maths exam where the dependent variable was the exam mark
(measured from 0 to 100), and the independent variables were revision time
(measured in hours) and intelligence (measured using IQ score). Here, it would be
possible to use an experimental design and manipulate the revision time of the
students. The tutor could divide the students into two groups, each made up of 50
students. In "group one", the tutor could ask the students not to do any revision.
Alternately, "group two" could be asked to do 20 hours of revision in the two weeks
prior to the test. The tutor could then compare the marks that the students achieved.
o Non-experimental research: In non-experimental research, the researcher does not
manipulate the independent variable(s). This is not to say that it is impossible to do
so, but it will either be impractical or unethical to do so. For example, a researcher
may be interested in the effect of illegal, recreational drug use (the independent
variable(s)) on certain types of behavior (the dependent variable(s)). However, whilst
possible, it would be unethical to ask individuals to take illegal drugs in order to study
what effect this had on certain behaviors. As such, a researcher could ask both drug
and non-drug users to complete a questionnaire that had been constructed to
indicate the extent to which they exhibited certain behaviors. Whilst it is not possible
to identify the cause and effect between the variables, we can still examine the
association or relationship between them. In addition to understanding the difference
between dependent and independent variables, and experimental and non-
experimental research, it is also important to understand the different characteristics
among variables.
9
What is measurement?
- The process of assigning numbers to observations.
Levels of Measurement
According to Stevens there are four levels of measurement:
1. Nominal Level (Categorical variables/discrete/qualitative)
Nominal variables are variables that have two or
more categories, but which do not have an intrinsic
order. For example, a real estate agent could
classify their types of property into distinct
categories such as houses, condos, co-ops or
bungalows. So "type of property" is a nominal
variable with 4 categories called houses, condos,
co-ops and bungalows. Of note, the different
categories of a nominal variable can also be
referred to as groups or levels of the nominal
variable. Another example of a nominal variable
would be classifying where people live in the USA
by state. In this case there will be many more
levels of the nominal variable (50 in fact).
Used as measures of identity, for qualitative
research.
Numbers indicate categories for purely
classification or identification purposes.
Numbers do not have values.
The categories are mutually exclusive (the
observations cannot fall into more than one
category)
The categories are exhaustive (there must be
enough categories for all the observations)
Examples:
Sex, religion, yes & no answers, political
parties, etc.
For sex:
M=1
F=2
Dichotomous variables are nominal variables
which have only two categories or levels. For
example, if we were looking at gender, we would
10
most probably categorize somebody as either
"male" or "female". This is an example of a
dichotomous variable (and also a nominal variable).
Another example might be if we asked a person if
they owned a mobile phone. Here, we may
categorize mobile phone ownership as either "Yes"
or "No". In the real estate agent example, if type of
property had been classified as either residential
or commercial then "type of property" would be a
dichotomous variable.
2. Ordinal Level (Categorical variables/discrete/qualitative)
Ordinal variables are variables that have two or
more categories just like nominal variables only
the categories can also be ordered or ranked. So if
you asked someone if they liked the policies of the
Democratic Party and they could answer either
"Not very much", "They are OK" or "Yes, a lot" then
you have an ordinal variable. Why? Because you
have 3 categories, namely "Not very much", "They
are OK" and "Yes, a lot" and you can rank them
from the most positive (Yes, a lot), to the middle
response (They are OK), to the least positive (Not
very much). However, whilst we can rank the
levels, we cannot place a "value" to them; we
cannot say that "They are OK" is twice as positive
as "Not very much" for example.
talks about how quality or which is better.
Used in measurement like ranking of individuals or
objects.
For qualitative data
Talks about how quality or which is better.
Numbers possess rank order characteristics.
Distances between any two successive numbers
are not necessarily equal
The categories must still be mutually exclusive and
exhaustive, but they also indicate the order of
magnitude of some variable.
Examples:
a. Likert’s scale
Strongly disagree = 1
11
Disagree = 2
Neither agree nor disagree = 3
Agree = 4
Strongly agree = 5
b. Student rankings in a class
3. Interval Level (Continuous variables/quantitative)
Interval variables are variables for which their
central characteristic is that they can be measured
along a continuum and they have a numerical
value (for example, temperature measured in
degrees Celsius or Fahrenheit). So the difference
between 20°C and 30°C is the same as 30°C to
40°C. However, temperature measured in degrees
Celsius or Fahrenheit is NOT a ratio variable.
Numbers that reflect differences among items.
Has all the properties of the ordinal scale
A given interval (distance) between scores has the
same meaning anywhere on the scale
Intervals provide information about how much
better one value is compared with another
Has no absolute zero
Examples:
scores in a test, grades, ages, blood
pressures,
temperature measured on 0Fahrenheit
&0Celsius
4. Ratio Level (Continuous variables/quantitative)
Ratio variables are interval variables, but with the
added condition that 0 (zero) of the measurement
indicates that there is none of that variable. So,
temperature measured in degrees Celsius or
Fahrenheit is not a ratio variable because 0°C
does not mean there is no temperature. However,
temperature measured in Kelvin is a ratio variable
as 0 Kelvin (often called absolute zero) indicates
that there is no temperature whatsoever. Other
examples of ratio variables include height, mass,
distance and many more. The name "ratio" reflects
the fact that you can use the ratio of
measurements. So, for example, a distance of ten
metres is twice the distance of 5 metres.
Highest type of scale.
12
Possesses all the characteristics of the interval
scale
Has a true or absolute zero point
The ratio of two values is meaningful
There can be no negative number.
Examples: measures of distance, length, height,
weight, time, cost of a car, loudness, width, etc.
Ambiguities in classifying a type of variable
In some cases, the measurement scale for data is ordinal, but the variable is
treated as continuous. For example, a Likert scale that contains five values - strongly
agree, agree, neither agree nor disagree, disagree, and strongly disagree - is ordinal.
However, where a Likert scale contains seven or more value - strongly agree,
moderately agree, agree, neither agree nor disagree, disagree, moderately disagree,
and strongly disagree - the underlying scale is sometimes treated as continuous
(although where you should do this is a cause of great dispute).
It is worth noting that how we categorize variables is somewhat of a choice.
Whilst we categorized gender as a dichotomous variable (you are either male or
female), social scientists may disagree with this, arguing that gender is a more
complex variable involving more than two distinctions, but also including
measurement levels like genderqueer, intersex and transgender. At the same time,
some researchers would argue that a Likert scale, even with seven values, should
never be treated as a continuous variable.
The Uses & Importance of Statistics
Sports
Education
Government
Healthcare / Medicines
Psychology
Sociology
Business & Economics
Other Fields
13
The Role of Statistics in Research
1. Helps the researcher in planning for the appropriate research
design.
2. Statistical tool can be used to determine the validity and reliability of
the data-gathering instrument.
3. Descriptive statistics are used to give meaning to the collected data.
4. Data-gathering involves getting information through interviews,
questionnaires, objective observations, experimentation,
psychological tests and other methods.
5. Statistics is used to test hypothesis.
6. Statistical analysis provides meaning and interpretation of the data.
Population Study Processes:
1. Collect
2. Organize
3. Present
4. Analyze
5. Conclusion
Sample Study Processes:
1. Sampling
2. Collect
3. Organize
4. Present
5. Analyze
6. Inferential
7. Conclusion
14
Statistical Symbols
Critical Value for Standard Normal
n Sample Size zα/2
Distribution
N Population Size z z-Test Statistic
∑ Sum df Degrees of Freedom
x¯ Sample Mean tα/2 Critical Value for t-Distribution
mu Population Mean t t Test Statistic
s Sample Standard Deviation D Difference
σ Population Standard Deviation D¯ Mean of the Differences
s2 Sample Variance sD Standard Deviation of the Differences
σ2 Population Variance F F Test Statistic
CV Coefficient of Variation Fα F Critical Value
z Z-score χ2 Chi-Square Test Statistic
IQR Interquartile Range χ2α Chi-Square Critical Value
Pi Percentile E Expected Value
Qi Quartile O Observed Value
P(x) Probability of X=x k Number of Groups
AC Complement Rule SS Sum of Squares
A∩B Intersection MS Mean Square
A∪B Union r Sample Correlation Coefficient
A|B Conditional Probability ρ Population Correlation Coefficient
n! Factorial Rule R Multiple Correlation Coefficient
nCr Combination Rule R2 Coefficient of Determination
nPr Permutation Rule 1−R2 Coefficient of Nondetermination
Predicted Value of the Regression
p Probability of a Success y^
Equation
q Probability of a Failure b0 Sample Y-Intercept Coefficient
p Population Proportion b1 Population Slope Coefficient
p P-value ei Residual
Significance Level = P(Type I Standard Error of Estimate =
α sest=s
Error) = False Positive Rate Standard Deviation of the Residuals
1−α P(True Negative) p Number of predictors in regression
P(Type II Error) = False
β ε Error term in regression
Negative Rate
1−β Power = P(True Positive) R Sum of the Ranks
H0 Null Hypothesis ws Wilcoxon Signed-Rank Test Statistic
H1=Ha Alternative Hypothesis U Mann-Whitney U Test Statistic
15
Gender Statistics and Indicators
Gender statistics refer to data that are collected, analyzed, and presented by sex,
and that reflect gender issues in society. They aim to reveal differences in the
conditions, needs, and opportunities of women and men, and to monitor progress
toward gender equality (United Nations, 2021). Unlike general statistics, gender
statistics are designed to capture disparities in areas such as education, health,
labor market participation, political representation, and access to resources.
Gender indicators are specific, measurable variables used to assess the status of
women and men in various aspects of life. These indicators can be quantitative—
such as female labor force participation rates—or qualitative—such as perceptions of
gender-based discrimination (Buvinic et al., 2020). They are often disaggregated by
sex and may also consider intersecting factors such as age, location, disability, and
socioeconomic status to provide a more comprehensive understanding of inequality
(OECD, 2022).
Examples of widely used gender indicators include:
Education: Literacy rates, school enrollment ratios, and completion rates by
sex.
Health: Maternal mortality rate, life expectancy, and reproductive health
access.
Economic participation: Gender wage gap, unemployment rate by sex, and
share of women in managerial positions.
Political empowerment: Proportion of seats held by women in national and
local legislatures.
Gender statistics and indicators are essential for designing evidence-based policies,
monitoring progress toward the Sustainable Development Goals (SDGs), particularly
SDG 5 on gender equality, and ensuring that no group is left behind in development
efforts (World Bank, 2021). When properly utilized, they help governments and
organizations allocate resources more equitably, address systemic inequalities, and
promote inclusive growth (Buvinic et al., 2020).
o Interactive Lectures and Discussions
o Writing Reflection Papers
o Identifying Quantitative & Qualitative Variables
o Evaluating Levels of Measurement
o Formative Quizzes and Feedback Sessions
16
RECOMMENDED LEARNING MATERIALS AND RESOURCES FOR SUPPLEMENTARY
READING
[Link]
[Link]
Module
FLEXIBLE TEACHING-LEARNING MODALITY
o Face-to-Face Learning
o Blended Learning
o Fully Online Learning
o Modular Distance Learning
o Synchronous and Asynchronous Sessions
o Project-Based Learning
ASSESSMENT TASK
o Quizzes and Short Tests
o Problem-Solving Exercises
o Oral Presentation
o Examinations
o Software Application
Ben-Zvi, D., & Garfield, J. (2020). Teaching and learning statistics: International
perspectives. Springer. [Link]
Bluman, A. G. (2021). Elementary statistics: A step-by-step approach (11th ed.).
McGraw-Hill Education.
Buvinic, M., Noe, L., & Swanson, E. (2020). Sex-disaggregated data: Why it matters
and how to produce it. Data2X. [Link]
Coursera. (2025). What is a Statistician? Duties, Pay, and How to Become One.
Highlights the role of statisticians in psychology, education, marketing, sports,
healthcare, and government. Coursera
17
Franklin, C., Kader, G., Mewborn, D., Moreno, J., Peck, R., Perry, M., & Scheaffer, R.
(2020). The statistical education of teachers. American Statistical Association.
GeeksforGeeks. (2025, July 15). Introduction of Statistics and its Types.
GeeksforGeeks. Retrieved [Month Day, Year], from
[Link]
[Link]
istics_(Webb)/zz%3A_Back_Matter/24%3A_Symbols
National Center for Biotechnology Information. (2021). Statistical advances in
epidemiology and public health. PMC. Highlights how biostatistics informs
health policy, clinical decisions, and disease prevention. PMC
OECD. (2022). Gender equality and well-being: OECD framework. Organisation for
Economic Co-operation and Development. [Link]
en
Ridgway, J. (2021). Statistics for citizenship: The role of statistical literacy in
democratic societies. Statistics Education Research Journal, 20(2), 1–14.
[Link]
U. S. National Institute of Standards and Technology (NIST). (2023). Engineering
Statistics Handbook: Glossary of Symbols. NIST.
Triola, M. F. (2020). Elementary statistics (13th ed.). Pearson.
United Nations. (2021). Gender statistics manual. UN Department of Economic and
Social Affairs. [Link]
Wikipedia contributors. (2025, updated). Sociology of sport. Explores sports as social
phenomena, strengthening ties between sociology, economics, and
psychology.
Wikipedia contributors. (2025, updated). Sports analytics. Explains how statistical
analysis supports performance evaluation, fan engagement, and strategic
decision-making in sports. Wikipedia
Wikipedia contributors. (2025, updated). Statistics. Covers the application of
statistics in business, government, social sciences, and natural sciences.
Wikipedia
World Bank. (2021). Gender data portal: Indicators and statistics. The World Bank.
18