Hypothesis Testing Study Guide
Hypothesis Testing Study Guide
I. INTRODUCTION
Job well done Louisian GEMs. We are near to the end of the line. Keep the spirit and
continue soaring high for the remaining weeks of this school year.
This week, you shall be given another lesson to study and enrichment activities to solve.
Attached to this module is the weekly Study and Assessment Guide
For these weeks of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the tasks
expected of you. GOOD LUCK!
Albert, J. R., Albacea, Z., Ayaay, M. J., David, I., & de Mesa, I. (2016)
Statistics and Probability. Commission on Higher Education. Manila,
Philippines
II. LEARNING CONTENT
At this point in time, maybe you are thinking of a good topic to study for Practical Research 1.
Researchers are interested in answering many types of questions. For example, a health student
might want to develop medicines for CoVid-19 virus. A physician might want to know whether the
Covid-19 vaccine will give adverse effects to Filipinos after the 2nd dose. An educator might wish
to see whether this online/distance learning is better than the face-to-face learning. An economist
might want to study the comeback of the economy to its normal status during this pandemic.
Automobile manufacturers are interested in determining whether seat belts will reduce the
severity of injuries caused by accidents. These types of questions can be addressed through
statistical hypothesis testing, which is a decision-making process for evaluating claims about a
population.
Recall that in your first module in this subject, we discussed the two types of Statistics. The
descriptive statistics was discussed in Week 2 until Week 8. For this final term, we will deal with
the other type of Statistics which is the Inferential Statistics.
In every end is a new beginning. We ended the midterm but don’t forget those lessons because
we will be needing those lessons for better understanding of the Final Term topics.
The only sure way of finding the truth or falsity of a hypothesis is by examining the entire
population but that is impossible to do, so we opted to use a sample for the purpose of drawing
conclusions. Using a sample, we can save us time, energy, money and effort.
Special consideration is given to the null hypothesis because it is the hypothesis to be tested as to
whether it should be rejected or not. The alternative hypothesis, from the word itself, will be the
choice if the null hypothesis were to be rejected.
In order to state the hypothesis correctly, the researcher must translate correctly the claim into
mathematical symbols. There are three possible sets of statistical hypotheses:
1. Ho: parameter = specific value This is two-tailed test.
Ha: parameter ≠ specific value
2. Ho: parameter = specific value This is a left-tailed test (one-tailed).
Ha: parameter < specific value
1. Ho: parameter = specific value This is a right-tailed test (one-tailed).
2. Ha: parameter > specific value
Examples: State the null and alternative hypothesis of the following problems;
1. The average age of lawyers is greater than 25.4 years old.
Ho: μ = 25.4
Ha: μ > 25.4
2. Filipino students achieved an average score of 353 points in Mathematical Literacy, which
was significantly lower than the OECD average of 489 points (Programme for International
Student Assessment [PISA], 2018).
Ho: μ = 489
Ha: μ < 489
3. The Philippines ranks 110th out of 139 countries in terms of mobile data speed, having an
average of 18.49 megabits per second (Mbps) (Ookla’s Speedtest Global Index, 2020).
Ho: μ = 18.49
Ha: μ ≠ 18.49
4. Kids and teens age 8 to 18 spend an average of more than seven hours a day looking at
screens (American Heart Association [AHA], 2018).
Ho: μ = 7
Ha: μ > 7
5. Filipino families earned Php 313, 000, on average (Philippine Statistics Authority [PSA],
2018).
Ho: μ = 313,000
Ha: μ ≠ 313,000
6. It added that the NCR logged an average of 1,025 new cases per day over the past seven
days covering February 28 to March 6 (OCTA Research, 2021)
Ho: μ = 1,025
Ha: μ ≠ 1,025
You should be able to master stating the statistical hypothesis because that is the first step in
hypothesis testing which will be your guide to make a good conclusion in the end.
In addition to that, there are four possible outcomes. In reality, the null hypothesis may or may
not be true. The decision to reject or not to reject is on the basis of the data obtained from the
sample of the population.
A Type I Error occurs if one rejects the null hypothesis when it is true. A Type II Error occurs if one
does not reject the null hypothesis when it is false.
The decision is made on the basis of probabilities. That is, if there is a large difference between
the value of the parameter obtained from the sample and the hypothesized parameter, the null
hypothesis is probably not true. The next question the researcher would ask is “How large a
difference is necessary to reject the null hypothesis?” Here is where the level of significance is
used.
Generally, statisticians agree on using arbitrary significance levels: 0.10, 0.05 and 0.01. That is, if
the null hypothesis is rejected, the probability of a Type I Error will be 10%, 5% or 1% and the
probability of a correct decision will be 90%, 95% or 99% depending on the level of significance is
used. It means that if the α = 0.05, there is a 5% chance of rejecting a true null hypothesis.
3rd Step: Identify the appropriate test-statistic and determine whether it is one-tailed or two-
tailed.
One way of determining the type of test used in hypothesis testing is based on how the alternative
hypothesis is formulated. A one-tailed test is used when the alternative hypothesis is directional
which means that the value of the means is greater than (>) or less than (<) the other measure. A
one-tailed test is a hypothesis test for which the rejection lies at only one tail of the distribution.
One tailed test is classified as left-tailed or right tailed. If the population mean (µ) is less than the
specified value of 𝜇0 , then it is a left tailed test for which the alternative hypothesis can be
expressed as µ < 𝜇0 . It is a right-tailed test if the population mean (µ) is greater than the specified
value of 𝜇0 for which the alternative hypothesis can be expressed as µ > 𝜇0 .
A two-tailed test is used when the alternative hypothesis is non-directional which means that the
values if two measures of the same kind are not equal. A two-tailed test has a not equal sign (≠) in
the alternative hypothesis. When the population mean (µ) is not equal to specified value of 𝜇0 ,
then alternative hypothesis can be expressed as µ ≠ 𝜇0 . A two-tailed test is a hypothesis for which
the rejection region lies on both ends of distribution, one on the left and one on the right.
There are two tests that can be used for means: the z-test and the t-test (one-sample, two-sample
and paired) and the chi square test for the standard deviation. You will also learn other statistical
tests like Analysis of Variance (ANOVA), Pearson r Correlation, and Linear Regression.
4th Step: Determine the critical value and the rejection region/s.
A critical value is selected from a table for the appropriate test. The critical value determines the
critical and noncritical regions. The critical region or the rejection region is the range of the values
of the test value that indicates that there is a significant difference and that the null hypothesis
should be rejected. The noncritical or the nonrejection region is the range of values of the test
value that indicates that the difference was probably due to chance and that the null hypothesis
should not be rejected.
The rejection region can be located on both sides with the nonrejection region in the middle, or it
can be on the left side or the right side of the nonrejection region. A test with two rejection regions
is called two-tailed test. In this test, the null hypothesis should be rejected when the test value is
in either of the two critical regions. A one-tailed test indicates that the null hypothesis should be
rejected when the test values is in the critical region on one side of the parameter. As a one-tailed
test is either right-tailed when the inequality in the alternative hypothesis is greater than (>) or
left-tailed when the inequality in less than (<). To illustrate that, study the illustration below:
If the test is two-tailed, the critical value will either be positive or negative. If the test is left-tailed,
the critical value will be negative. If the test is right-tailed, the critical value will be positive.
Next, the researcher must perform the required test to compute the test statistic. A statistical test
uses the data obtained from a sample to make a decision about whether the null hypothesis should
be rejected. The numerical value obtained from a statistical test is called the test value.
Based on the test results, the researcher will reach a conclusion about the population under study.
Generalization:
To summarize, here are the steps in Hypothesis testing:
1. State the null and alternative hypothesis.
2. Select the level of significance.
3. Identify the test-statistic.
4. Determine the critical value and the rejection region/s.
5. Compute the test statistic.
6. State the decision rule.
7. State the conclusion.
The steps will be used for the succeeding lessons based on the different statistical tests.
Note: If you have questions regarding your lessons, feel free to message your subject teacher.
Final Term – Second Semester Week 2: April 5 – 9, 2021
I. INTRODUCTION
Hello, Louisian GEM! I hope you are doing great as we unravel the facts and some pieces
of important information of this week’s corresponding learning. Also, keep the spirit and continue
soaring high for the remaining weeks of this school year.
This week, you shall be given another lesson to study and learning task to accomplish.
Attached to this module is the weekly Study and Assessment Guide
For this week of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the tasks
expected of you. GOOD LUCK!
GENERAL MATHEMATICS P a g e 1 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Icutan, S. L. et, al. (2012). Statistics with Probability a
Comprehensive Approach. Jimczyville Publications. Malabon City,
Philippines
In this chapter, we will discuss the different tests of hypothesis and some information you need
to help you decide on the most appropriate hypothesis testing for your research later on.
Before we proceed to our main topic, ponder on the question, “Why do you think it is essential to
test and prove our claims?”
Simply, our claim will merely be an opinion and will not be considered as factual unless supported
by an evidence. Just like in Mathematics, an idea cannot be considered accurate unless it is backed
up or supported by an evidence. And in statistics, in order for our claims to be true and valid, it
should undergo first into hypothesis testing which will also minimize our chance of error.
In deciding for the best Test statistic that you can use in your study, you should know first the two
types of Test. We have the parametric and the non-parametric.
Nonparametric Test
The nonparametric test is defined as the hypothesis test which is not based on underlying
assumptions, i.e. it does not require population’s distribution to be denoted by specific
parameters.
The test is mainly based on differences in medians. Hence, it is alternately known as the
distribution-free test. The test assumes that the variables are measured on a nominal or
ordinal level. It is used when the independent variables are non-metric.
For now, we focus first on the simplest parametric tests which is the One Sample t-test and z-test.
GENERAL MATHEMATICS P a g e 2 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
One sample z-test and t-test
One-sample test is a test of hypothesis based on data contained in a single sample using 2 test
statistics known as the z-test and the t-test.
Stat Trivia!
In the year 1908, an Irish brewery employee named W.S. Gosset published a paper containing his derivation
of the equation of the probability distribution? During his time his employer does not allow any publication
of research any members of the company. As a remedy, the paper was secretly published under the name
“Student”, and thus, the distribution was so then was called the Student t-distribution, or now simply called
t-distribution. Under this distribution, Gosset assumed that the samples were selected from a normally
distributed population. But even with this restriction, it can be shown that sampling distribution of T for
samples from non-normally distributed population can still approximate the t-distribution provided that
the distribution is bell shaped.
Before we proceed to answering some of the sample problems, we recall first the steps in
Hypothesis testing which have been discussed during the week 1(final term) of our
correspondence learning.
GENERAL MATHEMATICS P a g e 3 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
- One way of determining the type of test used in hypothesis testing is based on how the
alternative hypothesis is formulated. A one-tailed test is used when the alternative
hypothesis is directional which means that the value of the means is greater than (>) or
less than (<) the other measure. A one-tailed test is a hypothesis test for which the rejection
lies at only one tail of the distribution. One tailed test is classified as left-tailed or right
tailed. If the population mean (µ) is less than the specified value of 𝜇0 , then it is a left tailed
test for which the alternative hypothesis can be expressed as µ < 𝜇0 . It is a right-tailed test
if the population mean (µ) is greater than the specified value of 𝜇0 for which the alternative
hypothesis can be expressed as µ > 𝜇0 .
- A two-tailed test is used when the alternative hypothesis is non-directional which means
that the values if two measures of the same kind are not equal. A two-tailed test has a not
equal sign (≠) in the alternative hypothesis. When the population mean (µ) is not equal to
specified value of 𝜇0 , then alternative hypothesis can be expressed as µ ≠ 𝜇0 . A two-tailed
test is a hypothesis for which the rejection region lies on both ends of distribution, one on
the left and one on the right.
-For the test statistic, use the one-sample z test is used when there is only one sample in
the experiment that is known, and both the standard deviation and the mean of the
population are known also when the sample size is at least 30. On the other hand, to
compare a sample mean with the population mean when the population standard
deviation is not known, and when the sample size is less than 30, the one-sample t-test is
used.
GENERAL MATHEMATICS P a g e 4 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
- In here, we use the t-distribution table when we compute the critical value for a t-test
and the z-score when we compute the critical value of z. Recall your week 7 lesson in
midterm on how to get the t-tabular value and the z-score. (The t-distribution and z-score
are attached below)
( 𝑥̄ − 𝝁)√𝒏 ( 𝑥̄ − 𝝁)√𝒏
𝒕𝒄𝒐𝒎𝒑𝒖𝒕𝒆𝒅 = 𝒛𝒄𝒐𝒎𝒑𝒖𝒕𝒆𝒅 =
𝒔 𝒔
Where, Where,
𝑥̄ = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑚𝑒𝑎𝑛 𝑥̄ = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑚𝑒𝑎𝑛
𝜇 = 𝑝𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑚𝑒𝑎𝑛 𝜇 = 𝑝𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑚𝑒𝑎𝑛
𝑛 = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑖𝑧𝑒 𝑛 = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑖𝑧𝑒
𝑠 = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑠 = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛
GENERAL MATHEMATICS P a g e 5 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
- If the computed value of the statistic is greater than the critical value, then reject the null
hypothesis.
-If the computed value of the statistic is less than the tabular or critical value, then do not
reject the null hypothesis.
7. Conclusion
- Based on the test results, the researcher will reach a conclusion about the population
under study.
Since the steps in Hypothesis testing have been comprehensively discussed, you are now ready to
face the actual battle field as to how you are going to prove your claims. Before we proceed, always
take note that there are always two possible outcomes; if the result confirms the hypothesis, then
you’ve made measurements. If the result is contrary to the hypothesis, then you’ve made a
discovery. We may have not proved our claims, but at least, we discovered something. Who knows
if this discovery will have a great change and impact to our self, our belief and eventually, the
society we live in.
Example 1.
A test of the breaking strengths of six rings manufactured by a company showed a mean breaking
strength of 7750 pounds and a standard deviation of 145 pounds, whereas the manufacturer
claimed a mean breaking strength of 8000 pounds. Can we support the manufacturer’s claim at
significant level of 0.05?
Solution:
Ho:
a. State the null and µ = 8000 pounds and the manufacturer’s claim is justified
alternative hypotheses. Ha:
µ < 8000 pounds and the manufacturer’s claim is not justified
Since it has been stated from the problem that the significant level
b. Select the level of is 0.05 therefore,
significance
α = 0.05
c. Identify the appropriate since n=6 and 6 is less than 30 and the value of the population
test-statistic and mean is less than the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = t-test (one-tailed)
Given, α = 0.05
GENERAL MATHEMATICS P a g e 6 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
d. Determine the critical 𝑑𝑓 = 𝑛 − 1 = 6 − 1 = 5
value and the rejection
region/s. Refer on the table α = 0.05,
𝑑𝑓 = 5, one-tailed test
Example 2.
A researcher believes that in recent years, women have been getting taller. She knows that 10
years ago. The average height of young adult women living in her city was 63 inches. The standard
deviation is unknown. She randomly selected eight young adult women currently residing in her
city and measures their heights.
The following data are obtained: [64, 66, 68, 60, 62, 65, 66, and 63].
Solution:
α = 0.05
c. Identify the appropriate since n=8 and 8 is less than 30 and the value of the population
test-statistic and mean is not equal to the specified value, therefore,
GENERAL MATHEMATICS P a g e 7 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
determine whether it is test statistic = t-test (two-tailed)
one-tailed or two-tailed.
d. Determine the critical We use the t-distribution table since our test-statistic is t-test.
value and the rejection
region/s. Given, α = 0.05
𝑑𝑓 = 𝑛 − 1 = 8 − 1 = 7
Critical value = ±2.365 (±, since it’s two tailed, the rejection
region is found at both ends of the normal curve)
e. Compute the test Solving for the mean and variance of the sample,
statistic. we have 𝑥̄ = 64.25 and 𝑠 2 = 6.5 then we get the square root
of the variance for us to solve the standard deviation.
Therefore, 𝑠 = 2.55
𝑥̄ = 64.25 𝑠 = 2.55 𝑛=8 𝜇 = 63
Substitute.
( 𝑥̄ − 𝜇)√𝑛 ( 64.25 − 63)√8
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = =
𝑠 2.55
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = 1.39
f. State the decision rule. Since t-computed (|1.39|) is less than the t-tabular value (|-
2.01|), we do not reject the null hypothesis (Ho) and reject
alternative hypothesis (Ha).
g. Conclusion There is no sufficient evidence that the women are getting taller
in the city.
Example 3.
According to the Department of Education, high school teachers work an average of 40 hours per
week during the school year. A district supervisor of a certain school surveyed 28 randomly
selected teachers and found out that they work an average of 42.6 hours a week and the standard
deviation was 3.75 hours. Test if the mean number of hours worked by teachers in the supervisor’s
school differs from the national average. Use α = 0.01
Solution:
a. State the null and Ho:
alternative hypotheses. µ = 40 hours and the hours worked by teachers in the supervisor’s
school does not differ from the national average
GENERAL MATHEMATICS P a g e 8 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Ha:
µ ≠ 40 hours and the hours worked by teachers in the supervisor’s
school differs from the national average
b. Select the level of Since it has been stated from the problem that the significant level
significance is 0.01 therefore,
α = 0.01
c. Identify the appropriate since n=28 and 28 is less than 30 and the value of the population
test-statistic and mean is not equal to the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = t-test (two-tailed)
d. Determine the critical We use the t-distribution table since our test-statistic is t-test.
value and the rejection
region/s. Given, α = 0.05
𝑑𝑓 = 𝑛 − 1
𝑑𝑓 = 28 − 1
𝑑𝑓 = 27
Critical value = ±2.771 (±, since it’s two tailed, the rejection region
is found at both ends of the normal curve)
f. State the decision rule. Since t-computed (|3.67|) is greater than the t-tabular value
(|±2.771 |), we reject the null hypothesis (Ho) and do not reject
alternative hypothesis (Ha).
GENERAL MATHEMATICS P a g e 9 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Example 4.
A manufacturer claims that the average life of batteries used in their electronic games is 150 hours.
It is known that the standard deviation of this type of battery is 20 hours. A consumer wishes to
test the manufacturer’s claim and accordingly tests 100 electronic games using this battery and
found out that the mean is equal to 144 hours. Use α=0.05.
Solution:
a. State the null and Ho:
alternative hypotheses. µ = 150 hours (The average life of batteries used in electronic
games is equal to 150 hours)
Ha:
µ < 150 hours (The average life of batteries used in electronic
games less than 150 hours)
b. Select the level of Since it has been stated from the problem that the significant level
significance is 0.05 therefore,
α = 0.05
c. Identify the appropriate since n=100 and 100 is more than 30 and the value of the
test-statistic and population mean is less than the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = z-test (one-tailed)
d. Determine the critical We use the table for critical values of z since our test-statistic is z-
value and the rejection test.
region/s.
Given, α = 0.05
f. State the decision rule. Since t-computed (|-3.00|) is greater than the z-critical value
(|±1.645 |), we reject the null hypothesis (Ho) and do not reject
alternative hypothesis (Ha).
g. Conclusion The evidence is not enough to prove that the life of batteries is
150 hours.
GENERAL MATHEMATICS P a g e 10 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Example 5
An achievement test was administered to thousands of pupils with mean score of 85 and a
standard deviation of 8. A random sample of 50 pupils were given the same test and showed an
average score of 83.20. Is there evidence to show that this group has lower performance than the
ones in general at 0.05 level of significance?
Solution:
a. State the null and Ho:
alternative hypotheses. µ = 85 (there is no evidence to show that the group has lower
performance than the ones in general)
Ha:
µ < 85 hours (the group has lower performance than the ones in
general)
b. Select the level of Since it has been stated from the problem that the significant level
significance is 0.05 therefore,
α = 0.05
c. Identify the appropriate since n=50 and 50 is more than 30 and the value of the population
test-statistic and mean is less than the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = z-test (one-tailed)
d. Determine the critical We use the table for critical values of z since our test-statistic is z-
value and the rejection test.
region/s.
Given, α = 0.05
f. State the decision rule. Since t-computed (|-1.59|) is less than the z-critical value (|±1.645
|), we do not reject the null hypothesis (Ho) and reject alternative
hypothesis (Ha).
GENERAL MATHEMATICS P a g e 11 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Example 6.
A store owner claimed that the average weight of a pack of crackers is 250 g with a standard
deviation of 20 g. Would you agree to this claim if a random sample of 50 packs of crackers showed
an average of 242 g, using a 0.05 level of significance?
Solution:
a. State the null and Ho:
alternative hypotheses. µ = 250 (each pack of crackers do not weigh 250 g on the average)
Ha:
µ ≠ 250 (each pack of crackers weigh 250 g on the average)
b. Select the level of Since it has been stated from the problem that the significant level
significance is 0.05 therefore,
α = 0.05
c. Identify the appropriate Since n=50 and 50 is more than 30 and the value of the population
test-statistic and mean is not equal to the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = z-test (two-tailed)
d. Determine the critical We use the table for critical values of z since our test-statistic is z-
value and the rejection test.
region/s.
Given, α = 0.05
f. State the decision rule. Since t-computed (|-2.82|) is greater than the z-critical value
(|±1.96 |), we reject the null hypothesis (Ho) and do not reject
alternative hypothesis (Ha).
g. Conclusion There is no evidence to show that each pack of crackers weigh 250 g.
GENERAL MATHEMATICS P a g e 12 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Generalizations.
Before we finally end this chapter, we summarize the ideas and the pieces of information with this
week’s lesson.
The z-test is a statistical hypothesis test that follows a normal distribution while T-test follows a
Student’s T-distribution.
The t-test is appropriate when you are handling small samples (n < 30) while z-test is appropriate
when you are handling moderate to large samples (n > 30). In decision making of both of the tests,
always remember that, if the computed value of the statistic is greater than the tabular or critical
value, reject the null hypothesis and accept the alternative hypothesis and if the computed value
of the statistic is less than the tabular or critical value, accept the null hypothesis and reject the
alternative hypothesis.
One thing that we can take away from this week’s lesson is that life is a hypothesis testing. You
don’t know the result but still you give it a shot because there’s a hope that you might land where
your results are. And in case it doesn’t work out, just the hypothesis of life.
GENERAL MATHEMATICS P a g e 13 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Finals – Second Semester Week 3: April 12-16, 2021
I. INTRODUCTION
This week, you shall be given another lesson to study and enrichment activities to solve.
Attached to this 8th week module is the weekly Study and Assessment Guide
For this week of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the tasks
expected of you. GOOD LUCK!
Myers, R., Walpole, R., [Link]. (2017). Probability & Statistics for
Engineers & Scientists, Pearson Education Limited, England.
Before we start our discussion let’s recall some of the basic terminologies that we need for today’s
discussion
Hypothesis Testing is a method that uses sample data to decide between two competing claims
about a population characteristic.
Hypothesis is a claim or statement either about the value of a single population characteristic or
about the values of several population.
The Null Hypothesis denoted by Ho is a claim about a population characteristic that is initially
assumed to be true.
The Alternative Hypothesis denoted by Ha is the competing claim.
Let’s recall the how to identify if what tailed-test we are going to apply using the given table.
Note: If the hypothesis is assuming that the two-sample means are equal, always use the two-
tailed test.
If the hypothesis is assuming that the one sample mean is greater than or less than the other
sample mean, always use the one-tailed test.
The z-test and t-test are two of the statistical tools which can be used compare or to study
two groups of data through the value of their means.
Z test
▪ One Sample Z-test is used when there is only one sample in the experiment that is
known, and both of the standard deviation and the mean of the population are known.
▪ Two Sample Z-Test is used when we want to compare the population means between
two independent sample groups.
For today’s discussion we will concentrate on Z-test specifically Two Sample Means.
̅
𝒙𝟏 − ̅
𝒙𝟐
𝒛=
(𝒔𝟏 )𝟐 (𝒔𝟐 )𝟐
√
𝒏𝟏 + 𝒏𝟐
𝑥̅1 → 𝑠𝑎𝑚𝑝𝑙𝑒 𝑚𝑒𝑎𝑛 𝑓𝑜𝑟 𝑔𝑟𝑜𝑢𝑝 1 𝑥̅2 → 𝑠𝑎𝑚𝑝𝑙𝑒 𝑚𝑒𝑎𝑛 𝑓𝑜𝑟 𝑔𝑟𝑜𝑢𝑝 1
𝑠1 → 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑓𝑜𝑟 𝑔𝑟𝑜𝑢𝑝 1 𝑠2 → 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑓𝑜𝑟 𝑔𝑟𝑜𝑢𝑝 2
𝑛1 → 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑠𝑢𝑏𝑗𝑒𝑐𝑡𝑠 𝑖𝑛 𝑔𝑟𝑜𝑢𝑝 1 𝑛2 → 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑠𝑢𝑏𝑗𝑒𝑐𝑡𝑠 𝑖𝑛 𝑔𝑟𝑜𝑢𝑝 1
The Tabular Value of Z or the z-critical value at Indicated Level of Significance (𝜶)
Test / 𝜶 0.005 0.01 0.05 0.10
One-Tailed ±2.58 ±2.33 ±1.645 ±1.28
Two-Tailed ±2.81 ±2.575 ±1.96 ±1.645
Neighborhood
Sample A Sample B
𝑥̅1 = 10,100 𝑥̅2 = 10,300
𝑠1 = 300 𝑠2 = 400
𝑛1 = 100 𝑛2 = 100
Solution for a:
a) The bank wishes to test the null hypothesis that the two neighborhoods have the same
mean income.
Note: Since the hypothesis is assuming that the two neighbors have the same mean income,
we will use a two-tailed sample.
1st Step: Formulate the null and alternative 𝜇1 = 𝑇ℎ𝑒 𝑚𝑒𝑎𝑛 𝑖𝑛𝑐𝑜𝑚𝑒 𝑜𝑓 𝑛𝑒𝑖𝑔ℎ𝑏𝑜𝑟 𝐴
hypothesis.
𝜇2 = 𝑇ℎ𝑒 𝑚𝑒𝑎𝑛 𝑖𝑛𝑐𝑜𝑚𝑒 𝑜𝑓 𝑛𝑒𝑖𝑔ℎ𝑏𝑜𝑟 𝐵
6th Step: State the conclusion The evidence is insufficient to prove that the
two neighborhoods have the same mean
income.
Solution for b:
Note: Since the hypothesis is assuming that the neighbor A has a greater mean income that
neighbor B we will use one-tailed test.
6th Step: State the conclusion The two neighborhoods do not have the
same mean income.
Example #2
Mr. Joe Rivera, a senior high school teacher claims that the students in his class will get a higher
score on the Statistics Examination than Ms. Marie Santiago’s class. The mean of Mr. Joe’s class
for 51 students is 23.3 and the standard deviation is 4.7 while the mean of Ms. Marie’s class for
47 students is 18.3 and the standard deviation of 5.3. Using α = 0.10 can the claim of Mr. Joe
be true?
1st Step: Formulate the null and alternative 𝜇1 = 𝑇ℎ𝑒 𝑠𝑐𝑜𝑟𝑒 𝑜𝑓 𝑀𝑟. 𝐽𝑜𝑒 ′ 𝑠 𝑐𝑙𝑎𝑠𝑠
hypothesis.
𝜇2 = 𝑇ℎ𝑒 𝑠𝑐𝑜𝑟𝑒 𝑜𝑓 𝑀𝑠. 𝑀𝑎𝑟𝑖𝑒 ′ 𝑠 𝑐𝑙𝑎𝑠𝑠
2nd Step: Specify the Level of Significance and Level of Significance: 0.10
decide whether a one-tailed test or two- Since the problem requires to use one-tailed
tailed test. sample test, therefore the tabular value or
the z-critical value for 0.10 using one-tailed
sample tests is ±1.28
3rd step: Decide the test statistic to be used. Z-test of Two Sample Means
Identify the all of the given. Mr. Joe’s Class Ms. Marie’s Class
𝑥̅1 = 23.3 𝑥̅2 = 18.3
𝑠1 = 4.7 𝑠2 = 5.3
𝑛1 = 51 𝑛2 = 47
Use the formula: 𝑥̅1 − 𝑥̅2
𝑧=
(𝑠1 )2 (𝑠2 )2
√
𝑛1 + 𝑛2
Substitute the given on the formula and 𝑥̅1 − 𝑥̅2
𝑧=
simplify (𝑠1 )2 (𝑠2)2
√
𝑛1 + 𝑛2
23.3 − 18.3
𝑧= → 𝑆𝑖𝑚𝑝𝑙𝑖𝑓𝑦
2 2
√(4.7) + (5.3)
51 47
5
𝑧=
√0.433 + 0.598
1
𝑧=
1.015
𝑧 = 0.985 → 𝐶𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒
5th Step: Make a decision. To make a decision we need to compare the
absolute value of computed value and the
absolute value of tabular value or the z-
critical value of the specified level of
significance used in the problem.
Example #3
The average weight of 50 Volleyball Players who are properly trained for 3 to 6 months was 76.8
kilogram with the standard deviation of 12.6 kilogram while the other 38 volleyball players who
were trained only for less than 3 months have an average weight of 72.5 kilogram with the
1st Step: Formulate the null and alternative 𝜇1 = The weight of volleyball players who
hypothesis. were properly trained for 3 to 6 months.
2nd Step: Specify the Level of Significance and Level of Significance: 0.05
decide whether a one-tailed test or two- Since the problem requires to use two-tailed
tailed test. sample test, therefore the tabular value or
the z-critical value for 0.05 using two-tailed
sample tests is ±1.96
3rd step: Decide the test statistic to be used. Z-test of Two Sample Means
th
4 Step: Compute for the value of the test statistics used using the sample data.
Example #4
The IQs of 32 students from the section A showed a mean of 107 with a standard deviation of
10. While the IQs of 41 students from the section B showed a mean of 112 with a standard
deviation of 8. Is there a significant difference between the IQs of St. Anselm and St. Augustine?
Use 0.01 level of significance.
2nd Step: Specify the Level of Significance and Level of Significance: 0.01
decide whether a one-tailed test or two- Since the problem requires to use two-tailed
tailed test. sample test, therefore the tabular value or
• If the number of samples is greater than 30 or at least 30 and the population standard
deviation is known, use Z test.
• Two Sample Z-Test is used when we want to compare the population means between
two groups.
• |𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒| < | 𝑧 − 𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 𝑣𝑎𝑙𝑢𝑒|, do not reject the null hypothesis (Ho).
• |𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒| > | 𝑧 − 𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 𝑣𝑎𝑙𝑢𝑒|, reject the null hypothesis (Ho).
I. INTRODUCTION
This week, you shall be given another lesson to study and enrichment activities to solve.
Attached to this 4th week module is the weekly Study and Assessment Guide.
For these weeks of the final term, the following shall be your guide for the different
lessons and tasks that you need to accomplish. Be patient, read it carefully before proceeding to
the tasks expected of you.
GOOD LUCK!
Myers, R., Walpole, R., [Link]. (2017). Probability & Statistics for
Engineers & Scientists, Pearson Education Limited, England.
In the previous week, the z test was used to test the difference between two means when the
population standard deviation was known and the variables were normally or approximately
normally distributed, or when both sample sizes were greater than or equal to 30. In many
situations, however, these conditions cannot be met-that is, the population standard deviations
are not known. In these cases, a t-test is used to test the difference between means when the
two sample are dependent and when the samples are taken from two normally or approximately
normally distributed populations. Samples are independent samples when they are not related.
Formula for the t-test for Testing the Difference Between Two Means-Independent Samples
Variances are assumed to be unequal:
(𝑥1 − 𝑥2 ) − (𝜇1 − 𝜇2 )
𝑡=
(𝑠12 ) (𝑠22 )
√ +
𝑛1 𝑛2
where:
𝑠1and 𝑠2, the sample standard deviations, are estimates of 𝜎1 and 𝜎2 , respectively.
𝑥1 and 𝑥2 are the sample means.
𝜇1 and 𝜇2 are the population means.
The formula
(𝒙𝟏 − 𝒙𝟐 ) − (𝝁𝟏 − 𝝁𝟐 )
𝒕=
(𝒔𝟐𝟏 ) (𝒔𝟐𝟐 )
√
𝒏𝟏 + 𝒏𝟐
Where 𝑥1 − 𝑥2 is the observed difference between sample means and where the expected value
𝜇1 − 𝜇2 is equal to zero when no difference between population mean is hypothesized. The
(𝑠2 ) (𝑠22)
denominator √ 𝑛1 + is the standard error of the difference between two means. Since
1 𝑛2
mathematical deviation of the standard error is somewhat complicated, it will be omitted here.
Example1:
The average size of a farm in Indiana Country, Pennsylvania is 191 acres. The average size of a
fam in Greene Country, Pennsylvania is 199 acres. Assume the data were obtained from two
Solution:
a. State the null and alternative hypothesis.
𝐻0: 𝜇1 = 𝜇2
𝐻𝑎: 𝜇1 ≠ 𝜇2
16 2.120
Example 2: A Statistics teacher wants to compare his two classes to see if the performed any
differently on Midterm examination. Class A had 25 students with an average score of 70,
standard deviation of 15. Class B had 20 students with an average score of 74, standard deviation
of 25. Using alpha 0.05, did theses two classes perform differently on the test?
Solution:
a. State the null and alternative hypothesis.
𝐻0: 𝜇𝐴 = 𝜇𝐵
𝐻𝑎: 𝜇𝐴 ≠ 𝜇𝐵
At 𝑎𝑙𝑝ℎ𝑎 = 0.05, can we conclude that the sizes of the mushrooms differ between the type of
soils used?
Solution:
a. State the null and alternative hypothesis.
𝐻0: 𝜇𝑡𝑦𝑝𝑒1 = 𝜇𝑡𝑦𝑝𝑒𝑏
𝐻𝑎: 𝜇𝑡𝑝𝑒1 ≠ 𝜇𝑡𝑦𝑝𝑒2
Girls 9 2 0.866
Assume that the data is normally distributed, Is there a difference in the mean amount of time
boys and girls aged three to eight study each day? Test at the 5% level of significance.
Solution:
Example 5: A US Magazine, Consumer Reports, carried out a survey of the calorie content of a
number of different brand of hotdogs. The calorie content of 20 beef and 17 poultry hotdogs was
recorded as below:
Is there a difference in calorie content between beef and poultry hotdogs? Use 5% level of
significance.
Solution:
Generalization
The two-sample t-test (also known as the independent samples t-test) is a method used to test
whether the unknown population means of two groups are equal or not. It is used if the following
conditions are met.
1. Data values are independent.
2. Data values are randomly sampled from two normal populations.
3. The two individual groups have equal variances.
I. INTRODUCTION
Hello, Louisian GEM! I hope you are doing great as we unravel the facts and some pieces
of important information of this week’s corresponding learning. Also, keep the spirit and continue
soaring high for the remaining weeks of this school year.
This week, you shall be given another lesson to study and learning task to accomplish.
Attached to this module is the weekly Study and Assessment Guide
For this week of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the tasks
expected of you. GOOD LUCK!
On the previous weeks, you have learned about hypothesis testing and some statistical tools to
be used. In this section, a different version of the t- test is explained. This version is used when the
samples are dependent. Samples are considered to be dependent samples when the subjects are
paired or matched in some way.
Did you observe that your teacher, sometimes conducts a pre-test then a post-test? Why do you
think he did that? Simply because he wants to determine if you have an improvement or your
performance on your pre-test differs from your post- test.
What do you think your teacher do to determine if you have improved on that certain two sets
of tests? Well, the appropriate test is a paired sample t-test.
The paired sample t-test, sometimes called the dependent sample t-test, is a statistical procedure
used to determine whether the mean difference between two sets of observations is zero. In a
paired sample t-test, each subject or entity is measured twice, resulting in pairs of observations.
Common applications of the paired sample t-test include case-control studies or repeated-
measures designs. Suppose you are interested in evaluating the effectiveness of a company
training program. One approach you might consider would be to measure the performance of a
sample of employees before and after completing the program, and analyze the differences using
a paired sample t-test.
For example, suppose a medical researcher wants to see whether a drug will affect the reaction
time of its users. To test this hypothesis, the researcher must pretest the subjects in the sample
first. That is, they are given a test to ascertain their normal reaction times. Then after taking the
drug, the subjects are tested again, using a post-test. Finally, the means of the two tests are
compared to see whether there is a difference. Since the same subjects are used in both cases the
• The null hypothesis (Ho) assumes that the true mean difference (μd) is equal to zero.
• The two-tailed alternative hypothesis (Ha) assumes that μd is not equal to zero.
• The upper-tailed alternative hypothesis (Ha) assumes that μd is greater than zero.
• The lower-tailed alternative hypothesis (Ha) assumes that μd is less than zero.
The mathematical representations of the null and alternative hypotheses are defined below:
Note. It is important to remember that hypotheses are never about data, they are about the
processes which produce the data. In the formulas above, the value of μd is unknown. The goal of
hypothesis testing is to determine the hypothesis (null or alternative) with which the data are
more consistent.
Before we proceed to answering some of the sample problems, we recall first the steps in
Hypothesis testing which have been discussed during the 1st week of your module.
- If the computed value of the statistic is greater than the critical value, then reject the null
hypothesis.
-If the computed value of the statistic is less than the tabular or critical value, then do not
reject the null hypothesis.
7. Conclusion
- Based on the test results, the researcher will reach a conclusion about the population
under study.
Example 1.
To determine whether the students’ performance in College Algebra will improve after enrolling
in the subject for one term, a 60-item pre-test and post-test are administered to them on the first
day and last day of classes respectively. Use 1% level of significance.
Solution:
Since it has been stated from the problem that the significant level
b. Select the level of is 1% therefore,
significance
α = 0.01
α = 0.05, 𝑑𝑓 = 𝑛 − 1 = 10 − 1 = 9
d. Determine the critical
value and the rejection Refer on the table α = 0.01,
region/s. 𝑑𝑓 = 9, one-tailed test
f. State the decision rule. Since the t-computed (3.16) is greater than the t-tabular value (-
2.821), we reject the null hypothesis (Ho).
Example 2.
A study was conducted to determine the effectiveness of a weight loss program. The table below
shows the before and after weight of 12 subjects in the program. Is this program effective for
reducing weight? (Use a 5% significance level
Solution:
a. State the null and Ho: The weight loss program is effective.
alternative hypotheses. 𝜇𝑑 ≥ 0
Ha: The weight loss program is not effective.
𝜇𝑑 < 0
b. Select the level of The significance level was not mentioned in the problem. In this
significance case, we use the 5%.
c. Identify the appropriate since n=12 and 12 is less than 30 and the value of the population
test-statistic and mean is not equal to the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = paired sample t-test (one - tailed)
d. Determine the critical We use the t-distribution table since our test-statistic is paired
value and the rejection sample t-test.
region/s.
Given, α = 0.05, 𝑑𝑓 = 𝑛 − 1 = 12 − 1 = 11
Therefore,
𝒔𝒅 = 𝟏𝟐. 𝟕𝟖
Σ𝑑 = 132 𝑛 = 12
Substitute.
Σ𝑑 132
− 𝜇𝑑
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = 𝑛 = 12 − 0 = 11 = 𝟐. 𝟗𝟖
𝑠𝑑 12.78 3.69
√𝑛 √12
f. State the decision rule. Since t-computed (2.98) is greater than the t-tabular value (1.796),
we reject the null hypothesis (Ho).
g. Conclusion There is enough evidence to prove that the weight loss program is
effective.
Subject 1 2 3 4 5 6
Before (x1) 210 235 208 190 172 244
After (x2) 190 170 210 188 173 228
Solution:
a. State the null and Ho: The person’s cholesterol level will not change if the diet is
alternative hypotheses. supplemented by a certain mineral.
𝜇𝑑 = 0
Ha: The person’s cholesterol level will change if the diet is
supplemented by a certain mineral.
𝜇𝑑 ≠ 0
b. Select the level of Since it has been stated from the problem that the significant level
significance is 0.01 therefore,
α = 0.01
c. Identify the appropriate since n=6 and 6 is less than 30 and the value of the population
test-statistic and mean is not equal to the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = paired sample t-test (two-tailed)
d. Determine the critical We use the t-distribution table since our test-statistic is paired
value and the rejection sample t-test.
region/s.
Given, α = 0.01
𝑑𝑓 = 𝑛 − 1 = 6 − 1 = 5
Refer on the table α = 0.01,
𝑑𝑓 = 5, two-tailed test
Critical value = ±4.032 (±, since it’s two tailed, the rejection region
is found at both ends of the normal curve)
𝑑 = 100 𝑛=6
Substitute.
Σ𝑑 100
− 𝜇 𝑑 16.67
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = 𝑛𝑠 = 6 = = 𝟏. 𝟔𝟎𝟕
𝑑 25.39 10.37
√𝑛 √6
f. State the decision rule. Since t-computed 1.607 is less than the t-tabular value ±4.032, we
do not reject the null hypothesis (Ho) .
Generalizations.
Before we finally end this chapter, we summarize the ideas and the pieces of information with this
week’s lesson.
The paired sample t- test is a statistical hypothesis used to test a difference between means for
dependent samples. In testing hypothesis, follow the steps in hypothesis testing. First, state the
hypothesis and identify the claim. Remember that paired sample t-test is a statistical procedure
used to determine whether the mean difference between two sets of observations is zero.
Therefore, your Ho would always be 𝜇𝑑 = 0 , which means that there is no difference or there is
no change between the two sets of test. Then, select the level of significance, identify the
appropriate test-statistic (paired sample t-test) and determine whether it is one-tailed or two-
For this week’s lesson, it helps us realize that in life, there is always a chance of improvement and
a place for change.
I. INTRODUCTION
Job well done Louisian GEMs. We are near to the end of the line. Keep the spirit and continue
soaring high for the remaining weeks of this school year.
This week, you shall be given another lesson to study and enrichment activities to solve. Attached
to this module is the weekly Study and Assessment Guide
For these weeks of the final term, the following shall be your guide for the different lessons and
tasks that you need to accomplish. Be patient, read it carefully before proceeding to the tasks expected of
you. GOOD LUCK!
Albert, J. R., Albacea, Z., Ayaay, M. J., David, I., & de Mesa, I. (2016)
Statistics and Probability. Commission on Higher Education. Manila,
Philippines
A researcher might, for example, test students from multiple colleges to see if students from one of the
colleges consistently outperform students from the other colleges.
What statistical test will he use in his study? Another thing, for example, we might ask whether the
difference between two sample means could have been produced by chance.
How would we compare all of those means? What should we do, run a lot of t-tests, comparing every
possible combination of means?
Analysis of Variance commonly known as ANOVA is a technique that uses the F test to test a hypothesis
concerning the means of three (3) or more populations.
The t- and z-test methods developed in the 20th century were used for statistical analysis until 1918, when
Sir Ronald Fisher created the analysis of variance method. ANOVA is also called the Fisher analysis of
variance, and it is the extension of the t- and z-tests. The term became well-known in 1925, after
appearing in Fisher's book, "Statistical Methods for Research Workers. It was employed in experimental
psychology and later expanded to subjects that were more complex.
Two different estimates of the population variance are made. The first estimate is called the between-
group variance. It involves computing the variance by using the means of the groups or between the
groups. The second estimate, the within-group variance, is made by computing the variance using all the
data. It is not affected by differences in the means.
For a test of difference among three or more means, the following hypotheses are used:
Ho: μ1 = μ2 = μ3 = … = μn
Ha: At least one mean is different from the others
If there is no difference in the means, the between – group variance estimate will approximately be equal
to the within-group variance estimate, and the F test value will be approximately equal to 1. When the
means differ significantly, the between-group variance will be much larger than the within-group variance.
Then the F test value will be significantly greater than 1. Then the null hypothesis will be rejected.
The degree of freedom (df) for the F test are d.f.N = k - 1 where k is the number of groups and
d.f.D = n - k where n is the sum of the sample sizes of the groups, n = n 1+n2+…+nk. The sample sizes need
not be equal. The F test to compare the means is always a right-tailed test.
D. Solve for the variance between samples (MSB) and the variance within samples (MSW).
𝑆𝑆𝐵 𝑆𝑆𝑊
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
𝑘−1 𝑛−𝑘
where k – 1 and n – k are respectively, the degrees of freedoms for the numerator and
denominator for the F distribution.
F. After solving all the needed computations, input all computations on the ANOVA Table
The sum of SSB and SSW is called the total sum of squares and it is denoted by SST. That is, SST = SSB
+SSW
When the null hypothesis is rejected in ANOVA, it only means that at least one mean is different from the
others. To locate the difference or differences among the means, it is necessary to use other tests such as
the Tukey or the Scheffé test.
Use the following table for the critical value depending on the level of significance, degree of freedom of
Numerator (d.f.N.) and degree of freedom of Denominator (d.f.D.)
1. An introductory Calculus course is taken by students with varying high school records. The values
given are the final numerical averages in the Calculus course. At α = 0.05, test the claim that the
mean scores are equal in the three groups.
A B C
90 80 60
86 70 60
88 61 55
93 52 62
80 73 50
65 70
83
Solution:
Use the steps of the hypothesis testing together with the steps of solving the F test.
d.f.N = k – 1 = 3 – 1 = 2 d.f.D = n – k = 18 – 3 = 15
Use the F distribution table. Consider the intersection of d.f.N (degree of freedom of
Numerator), d.f.D (degree of freedom of Denominator) at the level of significance which
0.05.
B. Solve for the between – samples sum of squares (SSB). We need first to get the total values per
sample and the total number of values per sample.
Total of A: 437; n is 5
Total of B: 484; n is 7
Total of C: 357, n is 6
2
𝑇12 𝑇22 𝑇𝑘2 (∑ 𝑥)
𝑆𝑆𝐵 = ( + +⋯+ )−
𝑛1 𝑛2 𝑛𝑘 𝑛
D. Solve for the variance between samples (MSB) and the variance within samples (MSW).
𝑆𝑆𝐵 𝑆𝑆𝑊
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
𝑘−1 𝑛−𝑘
2162.44 1025.56
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
3−1 18 − 3
𝑀𝑆𝐵 = 1081.22 𝑀𝑆𝑊 = 68.37
2. A study was done before the recent surge in gasoline prices then to compare the cost to drive
25 miles for different types of hybrid vehicles. The cost of a gallon of gas at the time of the
study was approximately $2.50. Based on the information given below for different models
of hybrid cars, trucks, and SUVs, is there sufficient evidence to conclude a difference in the
mean cost to drive 25 miles? Use α = 0.05.
Hybrid Cars Hybrid SUVs Hybrid Trucks
2.10 2.10 3.62
2.70 2.42 3.43
1.67 2.25
1.67 2.10
1.30 2.25
d.f.N = k – 1 = 3 – 1 = 2 d.f.D = n – k = 12 – 3 = 9
Use the F distribution table. Consider the intersection of d.f.N (degree of freedom of
Numerator), d.f.D (degree of freedom of Denominator) at the level of significance which
0.05.
A. Compute the ∑ 𝑥 2 , square all values include in the different samples and then add.
∑ 𝑥 2 = 2.102 + 2.702 + 1.672 + 1.672 + 1.302 + 2.102 + 2.422 + 2.252 + 2.102 + 2.252
+ 3.622 + 3.432
∑ 𝑥 2 = 68.64
B. Solve for the between – samples sum of squares (SSB). We need first to get the total values per
sample and the total number of values per sample.
Total of Cars: 9.44; n is 5
Total of SUVs: 11.12; n is 5
Total of Trucks: 7.05, n is 2
2
𝑇12 𝑇22 𝑇𝑘2 (∑ 𝑥)
𝑆𝑆𝐵 = ( + +⋯+ )−
𝑛1 𝑛2 𝑛𝑘 𝑛
D. Solve for the variance between samples (MSB) and the variance within samples (MSW).
𝑆𝑆𝐵 𝑆𝑆𝑊
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
𝑘−1 𝑛−𝑘
3.88 1.24
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
3−1 12 − 3
𝑀𝑆𝐵 = 1.94 𝑀𝑆𝑊 = 0.14
3. The manager of a bank decided to compare the speed on the job of four tellers working in the
bank. The following data give the amount of time (in minutes) that they spent serving their
customers, picked at random.
Solution:
Use the steps of the hypothesis testing together with the steps of solving the F test.
d.f.N = k – 1 = 4 – 1 = 3 d.f.D = n – k = 20 – 4 = 16
Use the F distribution table. Consider the intersection of d.f.N (degree of freedom of
Numerator), d.f.D (degree of freedom of Denominator) at the level of significance which
0.025.
B. Solve for the between – samples sum of squares (SSB). We need first to get the total values per
sample and the total number of values per sample.
Teller 1: 22.2; n is 5
Teller 2: 22.1; n is 4
Teller 3: 25.7, n is 5
Teller 4: 23.7, n is 6
2
𝑇12 𝑇22 𝑇𝑘2 (∑ 𝑥)
𝑆𝑆𝐵 = ( + +⋯+ )−
𝑛1 𝑛2 𝑛𝑘 𝑛
D. Solve for the variance between samples (MSB) and the variance within samples (MSW).
𝑆𝑆𝐵 𝑆𝑆𝑊
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
𝑘−1 𝑛−𝑘
7.40 182.11
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
4−1 20 − 4
𝑀𝑆𝐵 = 2.47 𝑀𝑆𝑊 = 11.38
Solution:
Use the steps of the hypothesis testing together with the steps of solving the F test.
d.f.N = k – 1 = 4 – 1 = 3 d.f.D = n – k = 25 – 4 = 21
Use the F distribution table. Consider the intersection of d.f.N (degree of freedom of
Numerator), d.f.D (degree of freedom of Denominator) at the level of significance which
0.025.
H. Solve for the between – samples sum of squares (SSB). We need first to get the total values per
sample and the total number of values per sample.
Drug A: 30; n is 6
Drug B: 43; n is 8
Drug C: 25, n is 5
Drug D: 33, n is 6
2
𝑇12 𝑇22 𝑇𝑘2 (∑ 𝑥)
𝑆𝑆𝐵 = ( + +⋯+ )−
𝑛1 𝑛2 𝑛𝑘 𝑛
J. Solve for the variance between samples (MSB) and the variance within samples (MSW).
𝑆𝑆𝐵 𝑆𝑆𝑊
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
𝑘−1 𝑛−𝑘
1.19 183.38
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
4−1 25 − 4
𝑀𝑆𝐵 = 0.40 𝑀𝑆𝑊 = 8.73
Generalization:
We can use F-test to test the difference of more than 2 means. Follow the procedures in getting the F-
value and the hypothesis testing so that you can give a sound conclusion to a problem or study.
Note: If you have questions regarding your lessons, feel free to message your subject teacher.
I. INTRODUCTION
This week, you shall be given another lesson to study and enrichment activities to solve.
Attached to this 7th week module is the weekly Study and Assessment Guide
For this week of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the
tasks expected of you. GOOD LUCK!
Myers, R., Walpole, R., [Link]. (2017). Probability & Statistics for
Engineers & Scientists, Pearson Education Limited, England.
I. LEARNING CONTENT
The main objective of many statistical investigation is to establish relationship between two
variables. For example: A health researcher wants to determine the relationship between a
mother’s weight and her baby’s birth weight, An educational researcher wants to correlate the
grade of his students in science and math, A psychologist may want to know the relationship
between the IQ and personality of Individual.
These are just some examples of the situations that we can apply correlational test.
➢ This test also attempts to measure the strength of such relationships between two
variables namely (x and y) by the means of single number called a correlation coefficient.
➢ The variables x and y are referred to as bivariate variable.
Among the wide classes of correlational test, The Pearson Product Moment Coefficient of
Correlation (Pearson-r) and the Spearman Rank-Order Coefficient of Correlation are the two
most commonly used.
Correlation Coefficient
A correlation coefficient is a number between -1 and +1 which provides a measure of
the strength or degree of the linear association between two variables, x and y. The + sign and –
sign indicate only the direction of the correlation.
Remember: + is for a positive correlation and - is for a negative correlation.
0.00 No correlation
For example, if the result is 0.90 it indicates that it has a very high positive correlation but if the
result is -0.90 it indicates that it has a very high negative correlation. The signs “ + and – “
indicate the direction.
Scatter Diagram (Scatterpoint Diagram)
➢ This diagram is used to illustrate the relationship between two variables
➢ It is a graphic visualization of the relationship between the bivariate data.
3rd Step: Square each variable ( X and Y ) then compute for its summation
Science
Student Math Scores
Scores XY X2 Y2
Number (X)
(Y)
1 18 20 360 324 400
2 16 18 288 256 324
3 11 12 132 121 144
4 15 17 255 225 289
5 15 15 225 225 225
6 11 14 154 121 196
7 11 12 132 121 144
8 13 14 182 169 196
9 8 10 80 64 100
10 9 13 117 81 169
11 13 12 156 169 144
12 7 9 63 49 81
N=12 ∑ 𝑋 = 147 ∑ 𝑌 = 166 ∑ 𝑋𝑌 = 2144 ∑ 𝑋 2 = 1925 ∑ 𝑌 2 = 2412
A. State the null and alternative Ho: There is no significant relationship between the
hypothesis test scores in Mathematics and Science.
Ha: There is a significant relationship between the
test scores in Mathematics and Science.
C. Locate for the critical values on Using α=0.05 and 𝑑𝑓 = 10, employing two-tailed
the t-Distribution table, two-tailed test the tabular value is 2.228.
Since it’s a two-tailed test, the rejection region is on
the left of -2.228 and right of 2.228.
10
𝑡 = (0.92) (√ ) = (0.92)(√65.10
0.1536
𝑡 = 7.42
E. Statistical decision. Since it’s a two-tailed test, the rejection region is on
the left of -2.228 and right of 2.228.
The computed value is found on the rejection
region (right of 2.228) because 7.42 > 2.228.
Therefore, we reject the null hypothesis (Ho)
Since we already know the steps in identifying all the given to solve for r, we can just complete
the table.
Reading Comprehension
Student Vocabulary Test (Y)
(X) 𝑋𝑌 𝑋2 𝑌2
1 3 11
2 7 12
3 8 19
4 9 5
5 8 17
6 9 3
7 7 15
8 10 2
9 9 15
10 5 8
11 7 12
12 6 4
13 5 10
14 8 13
Compute for r
𝑟 = −0.02
d. State the conclusion Based on the quantitative interpretation of the degree of
the linear relationship, -0.02 is on the interval
-0.01 to -0.30 therefore it indicates that there is a
negligible negative correlation the reading
comprehension and the vocabulary of the 15 Grade 11
students of IV-Del Pilar.
A. State the null and alternative Ho: There is no significant relationship between the test
hypothesis scores in reading comprehension and the vocabulary of
the 15 Grade 11 students of St. Francis
Ha: There is a significant relationship between the test
scores in reading comprehension and the vocabulary of
the 15 Grade 11 students of St. Francis
13
𝑡 = (−0.02) (√ ) = (−0.02)(3.61)
0.996
𝑡 = −0.072
E. Make a decision. Since it’s a two-tailed test, the rejection region is on the
left of -2.16 and right of 2.16.
The computed value is found on the nonrejection region
(between -2.16 and 2.16). Therefore, we do not reject
the null hypothesis (Ho)
Compute for r
a. Identify all the given Given:
𝑛=9 ∑ 𝑋 = 36 ∑ 𝑌 = 370
𝑟 = 0.067
d. State the conclusion Based on the quantitative interpretation of the degree of
the linear relationship, 0.067 is between - 0.01 to - 0.30
therefore it indicates that there is a negligible positive
correlation between the number of hours a person
exercise and the amount of milk (in ounces) each person
consumes per week. Since the result is positive
A. State the null and alternative Ho: There is no significant relationship between the
hypothesis number of hours of exercise and the amount of milk
(in ounces) consumed per week.
Ha: There is a significant relationship between the
number of hours of exercise and the amount of milk
(in ounces) consumed per week.
7
𝑡 = (0.067)√
0.995
In studying relationships between two variables, collect the data and then construct a scatter
plot. The purpose of the scatter plot, as indicated previously, is to determine the nature of the
relationship. The possibilities include a positive linear relationship, a negative linear
relationship, a curvilinear relationship, or no discernible relationship. After the scatter plot is
drawn, the next steps are to compute the value of the correlation coefficient and to test the
significance of the relationship. If the value of the correlation coefficient is significant, the next
step is to determine the equation of the regression line, which is the data’s line of best fit.
(Note: Determining the regression line when r is not significant and then making predictions
using the regression line are meaningless.) The purpose of the regression line is to enable the
researcher to see the trend and make predictions on the basis of the data.
Figure 1 shows a scatter plot for the data of two variables. It shows that several lines can be
drawn on the graph near the points. Given a scatter plot, you must be able to draw the line of
best fit. Best fit means that the sum of the squares of the vertical distances from each point to
the line is at a minimum. The reason you need a line of best fit is that the values of y will be
predicted from the values of x; hence, the closer the points are to the line, the better the fit and
the prediction will be. See Figure 1. When r is positive, the line slopes upward and to the right.
When r is negative, the line slopes downward from left to right.
Figure 1
𝒚′ = 𝒂 + 𝒃𝒙
Example1: A statistic teacher recorded the Math and Science test scores of twelve senior high
students from one section. Find the equation of the regression.
Solution:
a. Given n=6
∑ 𝑋 = 147 , ∑ 𝑌 = 166, ∑ 𝑋𝑌 = 2144 , ∑ 𝑋 2 = 1925
b. Compute for a (∑ 𝑦)(∑ 𝑥 2 ) − (∑ 𝑥)(∑ 𝑥𝑦)
𝑎= 2
𝑛(∑ 𝑥 2 ) − (∑ 𝑥)
(166)(1925) − (147)(2144)
𝑎=
6(1925) − (147)2
𝒂 = −𝟎. 𝟒𝟑𝟒
c. Compute for b 𝑛(∑ 𝑥𝑦) − (∑ 𝑥)(∑ 𝑦)
𝑏= 2
𝑛(∑ 𝑥 2 ) − (∑ 𝑥)
6(2144) − (147)(166)
𝑏=
6(1925) − (1472 )
𝒃 = 𝟏. 𝟏𝟒𝟕
d. Substitute values of a and b 𝒚′ = 𝑎 + 𝑏𝒙
to the formula 𝒚′ = −𝟎. 𝟒𝟑𝟒 + 𝟏. 𝟏𝟒𝟕𝒙
Hence, the equation of the regression line 𝑦 ′ = 𝑎 + 𝑏𝑥 is 𝒚′ = −𝟎. 𝟒𝟑𝟒 + 𝟏. 𝟏𝟒𝟕𝒙
GENERALIZATION:
It is important to note that even if the correlation between two variables is high, it does not
necessarily mean causation. There are other possibilities, such as lurking variables or just a
I. INTRODUCTION
Hello, Louisian GEM! Welcome to the final week of our Correspondence Learning! Brace
yourself as we are about to reach the finish line!
This week, you shall be given one, last lesson to study. Attached to this module is the
Weekly Study and Assessment Guide
For this week of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the
tasks expected of you. GOOD LUCK!
We have already covered the different parametric tests in the past weeks of our
correspondence learning. We are now going to consider one of the most widely used non-
parametric test in the world of statistics which is the Chi-Square Test.
Most discussions of statistical procedures have been focused on the analysis of “quantitative”
data, but not all the variables in research are quantitative. Instead, they are “qualitative” in
nature.
Recall that: A qualitative data is one where the possible measurements are not measurements
but quantities or frequencies of things that occur in categories. Religious affiliation, gender, or
whichever brand of dishwashing detergent you prefer are examples.
A very useful test of this type is the Chi-Square test. It is used with ordinal data in the form of
frequencies or proportions. It uses counts or frequencies as data rather than means and
standard deviations. It requires the use of numerical values.
The 2 x 2 contingency table is one of the most common ways to summarize two qualitative
variables. To illustrate it, consider a sample of COVID-19 cases from Community A that was
classified according to symptoms and their sex.
Symptoms
Sex Total
Asymptomatic Symptomatic
Male 17 14 31
Female 10 9 19
Total 27 23 50
(𝑂 − 𝐸 ) 2 where:
𝑥2 = 𝛴 𝑂 = 𝑜𝑏𝑠𝑒𝑟𝑣𝑒𝑑 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦
𝐸
𝐸 = 𝑒𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦
As for the steps in hypothesis testing involving chi-square test, we follow the same steps in
hypothesis testing just like our previous lessons for the past weeks.
We recall the steps in hypothesis testing:
a. State the null and alternative hypotheses.
b. Select the level of significance.
c. Identify the appropriate test-statistic and determine whether it is one-tailed or two-tailed.
*note that chi-square test is always a right-tailed test since the critical values in the chi-square distribution are
all positive.
d. Determine the critical value and the rejection region.
*The chi-square distribution is being used. (The chi-square distribution table is found on the next page).
e. Compute the test statistic (use chi-square test formula)
f. State the decision rule
Reject the null hypothesis (Ho) if the computed value of the statistic is greater than the critical value.
Otherwise, do no reject the null hypothesis.
g. Conclusion
- Based on the test results, the researcher will reach a conclusion about the population under study.
The chi-square test of goodness-of-fitt is used if the researcher is interested in determining the
number of objects or responses which fall in different categories for a single qualitative variable
and wish to know if sample under analysis was drawn from a proportion with some specified
distribution or not.
Example 1:
The data for 150 students is recorded in the table below (the observed frequencies). We have
also indicated the expected frequency for each category. Since there are 150 measures or
observations and there are three categories (Dell, Asus, and Acer) we would indicate the
expected frequency for each category (same for all level) to be 150/3 or 50. Use 𝛼 = 0.05
significance level.
We now have the information we need to complete the steps in testing our statistical
hypotheses.
a. State the null and alternative Ho: There is no significant difference between the
hypotheses observed and the expected frequencies (we can
also say that the respondents show no preference)
𝛼 = 0.05, 𝑑𝑓 = 𝑐 − 1 = 3 − 1 = 2
d. Determine the critical value and
(3 categories → Dell, Asus, Acer)
Dell 27 50 10.58
Asus 59 50 1.62
Acer 64 50 3.92
2
(𝑂 − 𝐸 ) 2
𝑥 = 𝛴
𝐸
( 27 − 50)2 (59 − 50)2 (64 − 50)2
𝑥2 = + +
50 50 50
𝑥 2 = 10.58 + 1.62 + 3.92 = 16.12
f. State the statistical decision. Since computed (16.12) is greater than the critical
value (5.991), we reject the null hypothesis (Ho).
g. Conclusion There is a significant difference between the
observed and expected frequencies. We can also
say that the respondents show preference when it
comes to the brand of laptop they are going to buy.
Example 2:
There are three gates at the certain university. The building maintenance supervisor would like
to know if the gates are equally utilized. As an experiment, 600 are observed as they enter the
school. The number of students using each gate is reported below. At 0.01 significance level, can
we conclude that there is a difference in the use of the three gates?
Gate Number of students
Gate 1 245
Gate 2 205
Gate 3 150
Total 600
STATrivia!
Did you know that chi-square goodness-of-fit test was developed by Karl Pearson? The purpose
of this test is to determine how well an observed set of data fits an expected data.
In this test, two variables are involved and one variable is tested for independence to the other
variable.
Now let us consider the case of the two-variable chi-square test, also known as the Test of
Independence.
Example 1:
For example, we may wish to know if there is an association between the employment status
and the type of school where the employees graduated from. Use 0.05 level of significance.
The formula for chi-square is the same as before only that the degrees of freedom of a two-
variable chi square statistic is:
All the needed data are already complete, we’re now going to test our hypothesis.
Solution:
a. State the null and alternative Ho: There is no association between the employment status and the type of
hypotheses school.
Ha: There is association between the employment status and the type of
school.
Employment
Not Hired 21 20.1 0.04 3 4.2 0.34 9 9.6 0.04 12 11.1 0.07 45
Total 67 14 32 37 150
(𝑂 − 𝐸 ) 2
𝑥2 = 𝛴 = 0.02 + 0.04 + 0.15 + 0.34 + 0.02 + 0.04 + 0.03 + 0.07 = 0.71
𝐸
Since the computed value (0.71) is less than the critical value (7.815), do not
f. State the decision rule reject the null hypothesis (Ho).
g. Conclusion There is no association between the employment status and the type of
school they graduated from.
Social Status
Factors Affecting General Happiness TOTAL
SINGLE MARRIED WIDOWED
Friends and Social Life 41 49 42 132
Job or Primary Activity 27 50 33 110
Health and Physical Condition 12 21 25 58
TOTAL 80 120 100 300
We now have the information we need to complete the steps in testing our statistical
hypotheses for our research problem.
Solution:
Ho: There is no association between the social status and the factors
a. State the null and alternative affecting the general happiness.
hypotheses Ha: There is an association between the social status and the factors
affecting the general happiness.
𝛼 = 0.05
b. Select the level of significance
e. State the decision rule Since the computed value of 𝑥 2 (5.35) is less than the critical value (9.488),
do not reject Ho.
There is no association between the social status and the factors affecting
f. Conclusion
the general happiness.
When there is a close agreement between the value of O and E, the chi-square is small and the null
hypothesis is not rejected. When there are large differences between O and E, the chi-square is
large and the null hypothesis is rejected.
Before we finally close this final chapter of our correspondence learning, we should always bear
in our mind the essence of why we need to learn and study the different hypothesis testing that
we have had this whole semester. All of these things will train our mind in critical thinking
which will lead us to a better and wiser decision making. With these, we are not just thought to
solve problems but to infer better and deeper understanding of the situations that are
presented in front of us. We are able to unravel facts that we thought impossible. We are able
correct misconceptions that we always commit. Also, we discover things that yet to be
discovered by ourselves. Who knows that this small discovery of ours will change our life
forever? Most especially, we are able to weigh things out and decide what to reject and do not
accept in our life
Generalization:
Chi-square test is used when your data is nominal or ordinal which falls under qualitative. It is
used to compare the differences of the sample frequencies with the expected frequencies. It
has two applications namely: test of goodness-of-fit and test of association or independence.
where:
2
(𝑂 − 𝐸 ) 2 𝑂 = 𝑜𝑏𝑠𝑒𝑟𝑣𝑒𝑑 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦
𝑥 =𝛴
𝐸 𝐸 = 𝑒𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦
I hope you’ve had a very meaningful school year! I pray for the success of your academic
journey. May you have a meaningful and beautiful life ahead of you.