INFERENTIAL
STATISTICS
Descriptive statistics summarizes the data with
the purpose of describing what occurred in the
sample.
Inferential statistics are calculated with the
purpose of generalizing the findings from a
sample to the entire population of interest.
Inferential statistics will deal primarily on two kinds of analyses:
a. Difference
b. Relationship
1. State the null Hypothesis (H0)
2. Select an appropriate alternative hypothesis (H1)
3. Choose the appropriate statistical test
4. Select the desired level of significance to be
Steps for used. (The most common level is 0.05 but 0.01 is
also widely used)
testing
5. Compute the calculated value and determine the
hypothesis critical test value
6. Make the decision. Reject the null hypothesis if
the calculated value is larger than the critical
value, otherwise “do not reject” the null.
7. Make the conclusion.
One- tailed Test Level of Significance
(directional) Type
< less than
0.01 0.025 0.05 0.10
> greater than
≤ less than or equal to One- tailed ± 2.33 ± 1.86 ± 1.645 ± 1.28
≥greater than or equal to Two- tailed ± 2.575 ± 2.33 ± 1.96 ± 1. 645
Two- tailed
(non-
directional)
= equal or “is”
Three types of
Alternative
Hypothesis
Used when the researcher seeks of an increase
Positive of one variable over the other.
Directional H1= A > B
Hypothesis
The performance of Section A in English is
(one-tailed) greater than Section B’ s
Used when the researcher seeks of a decrease
of one variable over the other.
H1= A < B
Negative The performance of Section A in English is
Directional weaker than Section B’ s
Hypothesis Section A performed significantly lesser than
(one-tailed) Section B
Note: However, it is more difficult to use the negative
hypothesis, since it will be using negative numbers. Thus, it
is better to revert it to positive directional.
Used when the researcher desires to
know the difference only, and is not
Non- interested whether there is an increase or
directional decrease of variables.
hypothesis H1: A≠ B
(two-tailed)
Ex. The performance of Section A is not
equal to Section B’s
Used when the researcher can assume
that the population values are normally
distributed, variances are equal, and data
PARAMETRIC are interval or ratio in scale
TESTS z- test
f- test
t- test
Is used in comparing the difference of
two means
z- test Is used when n≥ 30
Is used when the standard deviation is
known
A sample mean compared to a population
mean
One sample
mean test.
/sample size
A sample mean with another mean.
Two sample
means (z-test)
The is a significant difference between the
performances of the two groups.
that the pupils learned better if they are grouped together or using the
cooperative learning.
When the sample size is less than 30
sample units
When the population standard deviation
is unknown
T- TEST
Steps in hypothesis testing Using the t-test.
Assuming the data are normally distributed, we
shall follow the steps and procedures enumerated
below:
Statement of the hypothesis
Choose the appropriate statistical test
Level of significance and find the degrees of freedom
(df) and the tabular or critical value using the table:
𝑑𝑓 = 𝑛 − 1 ( for one-sample mean test)
𝑑𝑓 = 𝑛1 + 𝑛2 − 2 ( for two-sample means test)
The degrees of freedom (df) are the number of
observations in the sample that free to vary around the
mean of the sample.
Computation
Decision
Conclusion
A consumer group is investigating a producer of diet
meals to examine if its prepackaged meals actually
contain the advertised 6 ounces of protein in each
package. Based on the following data, is there any
evidence that the meals do not contain the advertised
amount of protein? The goal is to test whether there is a
significant difference in the mean protein content and
the company’s claim. Run the appropriate test at a 5%
ONE SAMPLE level of significance
MEAN TEST 𝑥 =?, 𝑠 =?, 𝑛 = 20 , 𝜇 = 6 𝑜𝑢𝑛𝑐𝑒𝑠
I. Statement of the Hypothesis
𝐻𝑜 : The average protein content in the prepackage
diet meal equals to the claimed mean of 6.0 ounces
of protein
𝜇=6
𝐻𝑎 : The average protein content in the prepackage
diet meal equals to the claimed mean of 6.0 ounces
of protein
𝜇≠6
II. Statistical Test
𝑡 − 𝑡𝑒𝑠𝑡, since 𝑛 < 30
(𝑥 −𝜇)( 𝑛)
𝑡= 𝜎
III. Level of significance and
critical value:
𝛼 = 0.05
𝑑𝑓 = 𝑛 − 1 = 20 − 1 = 19
𝐶𝑉 = ±_________
[Link]
[Link]
IV. Computation
Given : 𝑛 = 20, 𝜇 = 6, 𝑠 = 0.55, 𝑥 = 5.4
(𝑥−𝜇)( 𝑛) (5.4−6)( 20) (−.6)(4.47) −2.68)
𝑡= 𝜎
= .55
= .55
= .55
= −4.88
𝑡 = −4.88
V. Decision
Since the computed value is −4.80 is greater than
− 2.093,therefore reject the 𝐻0 .
VI. Conclusion
The claim that that the protein content in the prepackaged meals is
6 ounces is not true.
Consider the previous example
I. Statement of hypothesis
𝐻𝑜 : There is no significant difference between the performances of
the two groups.
𝜇1 = 𝜇2
T-test using 𝐻𝑎 : There is significant difference between the performances of
the two groups.
the two- 𝜇1 ≠ 𝜇2
sample Means
II. Statistical Test
𝑡 − 𝑡𝑒𝑠𝑡, since 𝑛 < 30
𝑥1 − 𝑥2
𝑡=
𝑠 21 𝑠 2 2
𝑛1 + 𝑛2
III. Level of significance and critical
value:
𝛼 = 0.05
𝑑𝑓 = 𝑛1 + 𝑛2 − 2 = 27 + 23 − 2 = 48
𝐶𝑉 = 0.05,48 = ±2.011
[Link]
Student Experimental Group Control Group
1 22 16
IV. Computation
2 28 23
3 29 12
𝑥1 − 𝑥2 23.67 − 18.70
4 28 21 𝑡= = = 4.41
5 20 20 𝑠2
1 𝑠2 17.31 14.50
+ 𝑛2
6 17 14 𝑛1 2 27 + 23
7 23 13
8 28 18
9 23 20
10 28 20
11 19 24
12 17 14 V. Decision
13 25 18
14 23 22 Since the computed value of t that is 4.41
15 26 22 that is greater than the critical value 2.011,
16 30 12
17 17 17
therefore reject the 𝐻𝑜
18 26 17 VI. Conclusion
19 19 20
There is significant difference between the
20 28 24
21 24 19 performances of the two groups.
22 27 24
23 19 20
24 22
25 22
26 29
27 20
23.66666667 18.69565217
17.30769231 14.49407115
n=27 n=23
𝑑
𝑡= 𝑠
𝑛
2 ( 𝑑)2
T-test for 𝑑 −
𝑛
𝑠= 𝑛−1
Dependent
Sample Where : 𝑑= difference between means
(pre-test/post- 𝑑2 = sum of the squared difference
test) 𝑑 = sum of the mean difference
𝑛 = number of cases
𝑠 = standard deviation
Prof Leonardo Gabuyo conducted a review in his Eng 102
class. He gave an examination before and after the
review and gathered the following data:
stude Score before Score After
nt review Review d d2
1 16 18 -2 4
2 8 12 -4 16
3 12 10 2 4
4 10 17 -7 49
5 20 18 2 4
6 17 20 -3 9
7 9 11 -2 4
8 10 9 1 1
9 18 17 1 1
10 19 20 -1 1
-13 93
I. Statement of hypothesis
𝐻𝑜 : There is no significant difference between the mean
score of the students before and after the review class.
𝜇1 = 𝜇2
𝐻𝑎 : There is significant difference between the mean
score of the students before and after the review class.
𝜇1 ≠ 𝜇2
II. Statistical Test
𝑡 − 𝑡𝑒𝑠𝑡 𝑓𝑜𝑟 𝑑𝑒𝑝𝑒𝑛𝑑𝑒𝑛𝑡 𝑠𝑎𝑚𝑝𝑙𝑒, since 𝑛 < 30
𝑑
𝑡= 𝑠
𝑛
2 ( 𝑑)2
𝑑 − 𝑛
𝑠= 𝑛−1
III. Level of significance and critical value:
𝛼 = 0.05
𝑑𝑓 = 𝑛 − 1 = 10 − 1 = 9
𝐶𝑉 = 0.05,9 = ±2.262
IV. Computation V. Decision
Since the computed value is 1.41 is lesser
𝑑 −13
𝑑= = = 1.3 than ±2.262, therefore accept the 𝐻0 .
𝑛 10
VI. Conclusion
There is no significant difference between
( 𝑑)2
𝑑2 − the mean score of the students before and
𝑠= 𝑛−1
𝑛
after the review class. It implies that the
review class is not effective.
(−13)2
93−
= 10
= 2.91
10−1
𝑑 1.3
𝑡= 𝑠 = 2.91 = 1.41
𝑛 10
Analysis of variance (ANOVA) is a statistical technique that is
Analysis of used to check if the means of two or more groups are
Variance/ significantly different from each other.
ANOVA ANOVA checks the impact of one or more factors by comparing
the means of different samples.
(F-test)
𝑴𝑺𝑩 𝑀𝑒𝑎𝑛 𝑆𝑞𝑢𝑎𝑟𝑒𝑠 𝐵𝑒𝑡𝑤𝑒𝑒𝑛
𝑭= =
𝑴𝑺𝑾 𝑀𝑒𝑎𝑛 𝑆𝑞𝑢𝑎𝑟𝑒𝑠 𝑊𝑖𝑡ℎ𝑖𝑛
𝑺𝑺𝑩
𝑴𝑺𝑩 = 𝒅𝒇𝒃
𝑺𝑺𝑾
𝑴𝑺𝑾 = 𝒅𝒇𝒘
𝒙𝟐 𝑪 ( 𝒙)𝟐
𝑺𝑺𝑩 = −
𝒏 𝑵
( 𝒙)𝟐
𝑺𝑺𝑻 = 𝒙𝟐 − 𝑵
Analysis of 𝑺𝑺𝑾 = 𝑺𝑺𝑻 − 𝑺𝑺𝑩
𝒅𝒇𝒃 = 𝒌 − 𝟏
Variance/ 𝒅𝒇𝒘 = 𝒏(𝒌 − 𝟏)
ANOVA Where:
𝑭 = 𝑨𝑵𝑶𝑽𝑨
(F-test) 𝑴𝑺𝑩 = 𝑀𝑒𝑎𝑛 𝑆𝑞𝑢𝑎𝑟𝑒 𝐵𝑒𝑡𝑤𝑒𝑒𝑛
𝑴𝑺𝑾 = 𝑀𝑒𝑎𝑛 𝑆𝑞𝑢𝑎𝑟𝑒 𝑊𝑖𝑡ℎ𝑖𝑛
𝑺𝑺𝑩 = 𝑠𝑢𝑚 𝑜𝑓 𝑡ℎ𝑒 𝑠𝑞𝑢𝑎𝑟𝑒𝑠 𝑏𝑒𝑡𝑤𝑒𝑒𝑛
𝑺𝑺𝑻 = 𝑡𝑜𝑡𝑎𝑙 𝑠𝑢𝑚 𝑜𝑓 𝑡ℎ𝑒 𝑠𝑞𝑢𝑎𝑟𝑒𝑠
𝑺𝑺𝑾 = 𝑠𝑢𝑚 𝑜𝑓 𝑡ℎ𝑒 𝑠𝑞𝑢𝑎𝑟𝑒𝑠 𝑤𝑖𝑡ℎ𝑖𝑛
𝒅𝒇𝒃 = 𝑑𝑒𝑔𝑟𝑒𝑒𝑠 𝑜𝑓 𝑓𝑟𝑒𝑒𝑑𝑜𝑚 𝑏𝑒𝑡𝑤𝑒𝑒𝑛
𝒅𝒇𝒘 = 𝑑𝑒𝑔𝑟𝑒𝑒𝑛 𝑜𝑓 𝑓𝑟𝑒𝑒𝑑𝑜𝑚 𝑤𝑖𝑡ℎ𝑖𝑛
Analysis of
Variance/
ANOVA
(F-test)
Example:
In a learning experiment, 10 students are randomly
assigned to each of four groups. Each group is asked to perform
a set of tasks after exposure to the experimental treatment. Do
the groups differ in task performance? Use the raw score
method to test this problem at 5% level of significance.
Analysis of Student Group 1 Group 2 Group 3 Group 4
1 20 19 18 16
Variance/ 2 18 18 18 16
ANOVA 3
4
17
17
18
17
15
14
15
15
(F-test) 6
5 15
14
15
14
13
12
14
12
7 12 13 12 12
8 11 13 11 10
9 10 10 10 8
10 10 8 9 5
I. Statement of hypothesis
𝐻𝑜 : There is no significant difference among the
performance of the groups.
𝜇1 = 𝜇2 = 𝜇3 = 𝜇4
𝐻𝑎 : There is a significant difference among the
performance of the groups.
𝜇1 ≠ 𝜇2 ≠ 𝜇3 ≠ 𝜇4
II. Statistical Test
One-way Anova / F ratio
𝑀𝑆𝐵
F=
𝑀𝑆𝑊
III. Level of significance and critical value:
𝛼 = 0.05
dfB= 𝑛 − 1 = 4 − 1 = 3
dfW= 4 10 − 1 = 36
𝐶𝑉 = ±2.92
IV. Computation
V. Decision
Since the computed value of F-value is 0.899 is less
than the critical value 2.92, therefore accept the 𝐻0 .
VI. Conclusion
There is no significant difference between the
performances of the students exposed in three
groups of learning approach.
Pearson Product Moment Correlation (Pearson R)
Correlation Biserial Correlation
techniques Tetrachloric correlation
Nonparametric statistics are most
commonly used for variables at the
Nonparametric nominal or ordinal level of measurement,
Tests which basically means that they are used
for variables that do not have a normal
distribution.
The most common non-parametric test used for determining
Chi-square significant difference and association of variables
It is used when N is 20 or above
Test (X2)
It is used in data which are nominal or ordinal
The Chi-square test of independence is a statistical hypothesis
test used to determine whether two categorical or nominal
variables are likely to be related or not.
The data can be displayed in a contingency table where each row represents
a category for one variable and each column represents a category for the
other variable.
For example, say a researcher wants to examine the relationship between
gender (male vs. female) and preference of learning modality (online vs.
Independence modular)
H0: The two variables (factors) are independent.
There is no significant relationship between the respondents’ gender and
preference of learning modality.
Ha: The two variables (factors) are dependent.
There is a significant relationship between the respondents’ gender and
preference of learning modality.
[Link]
[Link]
[Link]
The test is applied to a single categorical variable from two or more different
populations. It is used to determine whether frequency counts are distributed
identically across different populations.
Use the test for homogeneity to decide if two populations with unknown
distributions have the same distribution as each other. In this case there will be a
single qualitative survey question or experiment given to two different
populations. The null and alternative hypotheses are:
Do male and female college students have the same distribution of living
conditions? Use a level of significance of 0.05. Suppose that 250 randomly
selected male college students and 300 randomly selected female college
students were asked about their living conditions: Dorm, Apartment, With
Homogeneity: Parents, Other.
H0: The two populations follow the same distribution.
There is no significant difference between the male and female college students’
living conditions.
Ha: The two populations have different distributions.
There is a significant difference between the male and female college students’
living conditions.
Use the goodness-of-fit test to decide whether a
population with an unknown distribution “fits” a known
distribution.
It is often used to evaluate whether sample data is
representative of the full population.
In this case there will be a single qualitative survey
Goodness-of- question or a single outcome of an experiment from a
Fit: single population.
The null and alternative hypotheses are:
H0: The population fits the given distribution.
Ha: The population does not fit the given distribution.
Use the test for independence to decide
whether two variables (factors) are
independent or dependent. In this case there
will be two qualitative survey questions or
experiments and a contingency table will be
constructed. The goal is to see if the two
Independence: variables are unrelated (independent) or
related (dependent). The null and alternative
hypotheses are:
H0: The two variables (factors) are independent.
Ha: The two variables (factors) are dependent.
Independence using Excel
[Link]
Spearman-Rho
Rank It is used when the two variables have ordinal data
Correlation
[Link]
7/pdf
[Link]
%[Link]