0% found this document useful (0 votes)
23 views100 pages

Hypothesis Testing Study Guide

The document outlines the final term schedule for students, focusing on hypothesis testing and its importance in statistical analysis. It provides a detailed guide on the steps involved in hypothesis testing, including formulating null and alternative hypotheses, selecting significance levels, and identifying test statistics. Additionally, it introduces the one-sample z-test and t-test, emphasizing the need for evidence to support claims in research.

Uploaded by

natashaadimaraa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
23 views100 pages

Hypothesis Testing Study Guide

The document outlines the final term schedule for students, focusing on hypothesis testing and its importance in statistical analysis. It provides a detailed guide on the steps involved in hypothesis testing, including formulating null and alternative hypotheses, selecting significance levels, and identifying test statistics. Additionally, it introduces the one-sample z-test and t-test, emphasizing the need for evidence to support claims in research.

Uploaded by

natashaadimaraa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Final Term – Second Semester Week 1: March 24-26, 2021

I. INTRODUCTION

Job well done Louisian GEMs. We are near to the end of the line. Keep the spirit and
continue soaring high for the remaining weeks of this school year.

This week, you shall be given another lesson to study and enrichment activities to solve.
Attached to this module is the weekly Study and Assessment Guide

DATE TOPIC ACTIVITIES OR TASKS


March 24-26, 2021 Hypothesis Testing Read the lessons.

For these weeks of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the tasks
expected of you. GOOD LUCK!

Content Hypothesis Testing


Learning Competencies The learner should be able to:
• illustrate a statistical hypothesis,
• differentiate a null hypothesis from alternative hypothesis,
• differentiate Type I from Type II error, and
• identify the steps in hypothesis tesing.
Activities Read the lessons.
Essential Questions How important is inference making in the different
disciplines?
Value Statement In every end is a new beginning.
References Textbooks:

Alferez, M. S. & Duro, M. C. A. (2018). Statistics and Probability.


MSA Publishing House. Cainta, Philippines.

De Guzman, D. (2017). Statistics and Probability. C & E Publishing


Inc. Quezon City, Philippines

Albert, J. R., Albacea, Z., Ayaay, M. J., David, I., & de Mesa, I. (2016)
Statistics and Probability. Commission on Higher Education. Manila,
Philippines
II. LEARNING CONTENT

At this point in time, maybe you are thinking of a good topic to study for Practical Research 1.
Researchers are interested in answering many types of questions. For example, a health student
might want to develop medicines for CoVid-19 virus. A physician might want to know whether the
Covid-19 vaccine will give adverse effects to Filipinos after the 2nd dose. An educator might wish
to see whether this online/distance learning is better than the face-to-face learning. An economist
might want to study the comeback of the economy to its normal status during this pandemic.
Automobile manufacturers are interested in determining whether seat belts will reduce the
severity of injuries caused by accidents. These types of questions can be addressed through
statistical hypothesis testing, which is a decision-making process for evaluating claims about a
population.

Recall that in your first module in this subject, we discussed the two types of Statistics. The
descriptive statistics was discussed in Week 2 until Week 8. For this final term, we will deal with
the other type of Statistics which is the Inferential Statistics.

Did you know?


Statistical inference as a distinct discipline began with Francis Galton (1822–1911), a cousin of
Charles Darwin, whose On the Origin of Species (1859) became the inspiration of Galton's life.

In every end is a new beginning. We ended the midterm but don’t forget those lessons because
we will be needing those lessons for better understanding of the Final Term topics.

1st Step: State the null and alternative hypotheses.


Remember that in every hypothesis testing begins with the statement of a hypothesis. A statistical
hypothesis is an inference about a population parameter. This inference may or may not be true.

The only sure way of finding the truth or falsity of a hypothesis is by examining the entire
population but that is impossible to do, so we opted to use a sample for the purpose of drawing
conclusions. Using a sample, we can save us time, energy, money and effort.

There are 2 Kinds of Hypothesis:


1. Alternative Hypothesis (Ha) states specific difference between a parameter and a specific
value.
2. Null Hypothesis (Ho) states that no difference between a parameter and a specific value.

Special consideration is given to the null hypothesis because it is the hypothesis to be tested as to
whether it should be rejected or not. The alternative hypothesis, from the word itself, will be the
choice if the null hypothesis were to be rejected.

In order to state the hypothesis correctly, the researcher must translate correctly the claim into
mathematical symbols. There are three possible sets of statistical hypotheses:
1. Ho: parameter = specific value This is two-tailed test.
Ha: parameter ≠ specific value
2. Ho: parameter = specific value This is a left-tailed test (one-tailed).
Ha: parameter < specific value
1. Ho: parameter = specific value This is a right-tailed test (one-tailed).
2. Ha: parameter > specific value

Examples: State the null and alternative hypothesis of the following problems;
1. The average age of lawyers is greater than 25.4 years old.
Ho: μ = 25.4
Ha: μ > 25.4
2. Filipino students achieved an average score of 353 points in Mathematical Literacy, which
was significantly lower than the OECD average of 489 points (Programme for International
Student Assessment [PISA], 2018).
Ho: μ = 489
Ha: μ < 489
3. The Philippines ranks 110th out of 139 countries in terms of mobile data speed, having an
average of 18.49 megabits per second (Mbps) (Ookla’s Speedtest Global Index, 2020).
Ho: μ = 18.49
Ha: μ ≠ 18.49
4. Kids and teens age 8 to 18 spend an average of more than seven hours a day looking at
screens (American Heart Association [AHA], 2018).
Ho: μ = 7
Ha: μ > 7
5. Filipino families earned Php 313, 000, on average (Philippine Statistics Authority [PSA],
2018).
Ho: μ = 313,000
Ha: μ ≠ 313,000
6. It added that the NCR logged an average of 1,025 new cases per day over the past seven
days covering February 28 to March 6 (OCTA Research, 2021)
Ho: μ = 1,025
Ha: μ ≠ 1,025

You should be able to master stating the statistical hypothesis because that is the first step in
hypothesis testing which will be your guide to make a good conclusion in the end.

In addition to that, there are four possible outcomes. In reality, the null hypothesis may or may
not be true. The decision to reject or not to reject is on the basis of the data obtained from the
sample of the population.

Reject Ho Do not Reject Ho


Ho is true Type I Error Correct decision
P=α
Ho is false Correct decision Type II Error
P=β

A Type I Error occurs if one rejects the null hypothesis when it is true. A Type II Error occurs if one
does not reject the null hypothesis when it is false.

The decision is made on the basis of probabilities. That is, if there is a large difference between
the value of the parameter obtained from the sample and the hypothesized parameter, the null
hypothesis is probably not true. The next question the researcher would ask is “How large a
difference is necessary to reject the null hypothesis?” Here is where the level of significance is
used.

2nd Step: Select the level of significance.


The level of significance, denoted by the Greek letter α (Alpha) is the maximum probability of
committing a Type I Error which is P(type I error) = α. The probability of Type II Error is denoted by
Greek letter β (Beta) which is P(type II error) = β.

Generally, statisticians agree on using arbitrary significance levels: 0.10, 0.05 and 0.01. That is, if
the null hypothesis is rejected, the probability of a Type I Error will be 10%, 5% or 1% and the
probability of a correct decision will be 90%, 95% or 99% depending on the level of significance is
used. It means that if the α = 0.05, there is a 5% chance of rejecting a true null hypothesis.

3rd Step: Identify the appropriate test-statistic and determine whether it is one-tailed or two-
tailed.

One way of determining the type of test used in hypothesis testing is based on how the alternative
hypothesis is formulated. A one-tailed test is used when the alternative hypothesis is directional
which means that the value of the means is greater than (>) or less than (<) the other measure. A
one-tailed test is a hypothesis test for which the rejection lies at only one tail of the distribution.
One tailed test is classified as left-tailed or right tailed. If the population mean (µ) is less than the
specified value of 𝜇0 , then it is a left tailed test for which the alternative hypothesis can be
expressed as µ < 𝜇0 . It is a right-tailed test if the population mean (µ) is greater than the specified
value of 𝜇0 for which the alternative hypothesis can be expressed as µ > 𝜇0 .

A two-tailed test is used when the alternative hypothesis is non-directional which means that the
values if two measures of the same kind are not equal. A two-tailed test has a not equal sign (≠) in
the alternative hypothesis. When the population mean (µ) is not equal to specified value of 𝜇0 ,
then alternative hypothesis can be expressed as µ ≠ 𝜇0 . A two-tailed test is a hypothesis for which
the rejection region lies on both ends of distribution, one on the left and one on the right.

There are two tests that can be used for means: the z-test and the t-test (one-sample, two-sample
and paired) and the chi square test for the standard deviation. You will also learn other statistical
tests like Analysis of Variance (ANOVA), Pearson r Correlation, and Linear Regression.
4th Step: Determine the critical value and the rejection region/s.

A critical value is selected from a table for the appropriate test. The critical value determines the
critical and noncritical regions. The critical region or the rejection region is the range of the values
of the test value that indicates that there is a significant difference and that the null hypothesis
should be rejected. The noncritical or the nonrejection region is the range of values of the test
value that indicates that the difference was probably due to chance and that the null hypothesis
should not be rejected.

The rejection region can be located on both sides with the nonrejection region in the middle, or it
can be on the left side or the right side of the nonrejection region. A test with two rejection regions
is called two-tailed test. In this test, the null hypothesis should be rejected when the test value is
in either of the two critical regions. A one-tailed test indicates that the null hypothesis should be
rejected when the test values is in the critical region on one side of the parameter. As a one-tailed
test is either right-tailed when the inequality in the alternative hypothesis is greater than (>) or
left-tailed when the inequality in less than (<). To illustrate that, study the illustration below:

If the test is two-tailed, the critical value will either be positive or negative. If the test is left-tailed,
the critical value will be negative. If the test is right-tailed, the critical value will be positive.

5th Step: Compute the test statistic.

Next, the researcher must perform the required test to compute the test statistic. A statistical test
uses the data obtained from a sample to make a decision about whether the null hypothesis should
be rejected. The numerical value obtained from a statistical test is called the test value.

6th Step: State the decision rule.


Based on the test results, the researcher will reach a conclusion about the population under study.
The two possible decisions are
(a) reject the null hypothesis, or
(b) do not reject the null hypothesis.
7th Step: State the conclusion.

Based on the test results, the researcher will reach a conclusion about the population under study.

Generalization:
To summarize, here are the steps in Hypothesis testing:
1. State the null and alternative hypothesis.
2. Select the level of significance.
3. Identify the test-statistic.
4. Determine the critical value and the rejection region/s.
5. Compute the test statistic.
6. State the decision rule.
7. State the conclusion.

The steps will be used for the succeeding lessons based on the different statistical tests.

Note: If you have questions regarding your lessons, feel free to message your subject teacher.
Final Term – Second Semester Week 2: April 5 – 9, 2021

I. INTRODUCTION

Hello, Louisian GEM! I hope you are doing great as we unravel the facts and some pieces
of important information of this week’s corresponding learning. Also, keep the spirit and continue
soaring high for the remaining weeks of this school year.

This week, you shall be given another lesson to study and learning task to accomplish.
Attached to this module is the weekly Study and Assessment Guide

DATE TOPIC ACTIVITIES OR TASKS


One sample z-test and t-test Read the lessons.
April 5 - 9, 2021
Answer the learning tasks

For this week of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the tasks
expected of you. GOOD LUCK!

Content One sample z-test and t-test


Learning Competencies The learner should be able to:
• illustrate the
(a) null hypothesis,
(b) alternative hypothesis,
(c) level of significance,
(d) rejection region; and
(e) types of errors in hypothesis testing
• differentiate t-test and z-test;
• identify the parameter to be tested given a real-life
problem; and,
• draw conclusion about the population mean based on the
test-statistic value and the rejection region.
Activities Written Assessment 1
Essential Questions Why do you need to test if your claims are true or not?
Value Statement There are always two possible outcomes; if the result confirms the
hypothesis, then you’ve made measurements. If the result is
contrary to the hypothesis, then you’ve made a discovery.
- Enrico Fermi
References Textbooks:

Altares, P. et, al. (2012). Elementary Statistics with Computer


Applications. Rex Boos store. Quezon City, Philippines.

GENERAL MATHEMATICS P a g e 1 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Icutan, S. L. et, al. (2012). Statistics with Probability a
Comprehensive Approach. Jimczyville Publications. Malabon City,
Philippines

Cabrera, R. et, al. (2014) Fundamental of Statistics. Grandbooks


Publishing. Manila, Philippines

II. LEARNING CONTENT

In this chapter, we will discuss the different tests of hypothesis and some information you need
to help you decide on the most appropriate hypothesis testing for your research later on.

Before we proceed to our main topic, ponder on the question, “Why do you think it is essential to
test and prove our claims?”
Simply, our claim will merely be an opinion and will not be considered as factual unless supported
by an evidence. Just like in Mathematics, an idea cannot be considered accurate unless it is backed
up or supported by an evidence. And in statistics, in order for our claims to be true and valid, it
should undergo first into hypothesis testing which will also minimize our chance of error.

In deciding for the best Test statistic that you can use in your study, you should know first the two
types of Test. We have the parametric and the non-parametric.

What is the difference between parametric test and Non-Parametric Test?


Parametric Test
The parametric test is the hypothesis test which provides generalizations for making
statements about the mean of the parent population. A t-test based on Student’s t-
statistic, which is often used in this regard.
The t-statistic rests on the underlying assumption that there is the normal distribution of
variable and the mean in known or assumed to be known. The population variance is
calculated for the sample. It is assumed that the variables of interest, in the population are
measured on an interval scale.

Nonparametric Test
The nonparametric test is defined as the hypothesis test which is not based on underlying
assumptions, i.e. it does not require population’s distribution to be denoted by specific
parameters.
The test is mainly based on differences in medians. Hence, it is alternately known as the
distribution-free test. The test assumes that the variables are measured on a nominal or
ordinal level. It is used when the independent variables are non-metric.

For now, we focus first on the simplest parametric tests which is the One Sample t-test and z-test.

GENERAL MATHEMATICS P a g e 2 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
One sample z-test and t-test

How do we define one-sample test?

One-sample test is a test of hypothesis based on data contained in a single sample using 2 test
statistics known as the z-test and the t-test.

What is the difference between the z-test and t-test?


The one-sample z test is used when there is only one sample in the experiment that is known, and both the
standard deviation and the mean of the population are known also when the sample size is at least 30. On
the other hand, to compare a sample mean with the population mean when the population standard
deviation is not known, and when the sample size is less than 30, the one-sample t-test is used.

Stat Trivia!
In the year 1908, an Irish brewery employee named W.S. Gosset published a paper containing his derivation
of the equation of the probability distribution? During his time his employer does not allow any publication
of research any members of the company. As a remedy, the paper was secretly published under the name
“Student”, and thus, the distribution was so then was called the Student t-distribution, or now simply called
t-distribution. Under this distribution, Gosset assumed that the samples were selected from a normally
distributed population. But even with this restriction, it can be shown that sampling distribution of T for
samples from non-normally distributed population can still approximate the t-distribution provided that
the distribution is bell shaped.

Before we proceed to answering some of the sample problems, we recall first the steps in
Hypothesis testing which have been discussed during the week 1(final term) of our
correspondence learning.

*Steps in Hypothesis Testing

1. State the null and alternative hypotheses.


-In any hypothesis testing problems, there are two competing hypotheses under
consideration. The null hypothesis denoted by Ho, is a claim about the population
characteristics that is initially assumed to be true. The alternative hypothesis denoted by
Ha is the competing claim.
2. Select the level of significance.
-The significance level which is usually denoted by alpha (α) is related to the degree of
certainty we require in order to reject the null hypothesis in favor of the alternative
hypothesis. The most common significance levels are 0.05 and 0.01 because of the desire
to maintain a low probability of rejecting the null hypothesis when it is in fact true. The use
of a level of significance (α) of 0.05 or 0.01 means that we are willing to commit an error
of 5% or 1% and are, therefore, confident of making 95% or 99% correct decision. If,
instance, we set the (α)=0.5, the probability of incorrectly rejecting the null hypothesis
when it is in fact true at 5% can be avoided or protected from error by choosing a lower
value for α.
3. Identify the appropriate test-statistic and determine whether it is one-tailed or two-tailed.

GENERAL MATHEMATICS P a g e 3 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
- One way of determining the type of test used in hypothesis testing is based on how the
alternative hypothesis is formulated. A one-tailed test is used when the alternative
hypothesis is directional which means that the value of the means is greater than (>) or
less than (<) the other measure. A one-tailed test is a hypothesis test for which the rejection
lies at only one tail of the distribution. One tailed test is classified as left-tailed or right
tailed. If the population mean (µ) is less than the specified value of 𝜇0 , then it is a left tailed
test for which the alternative hypothesis can be expressed as µ < 𝜇0 . It is a right-tailed test
if the population mean (µ) is greater than the specified value of 𝜇0 for which the alternative
hypothesis can be expressed as µ > 𝜇0 .
- A two-tailed test is used when the alternative hypothesis is non-directional which means
that the values if two measures of the same kind are not equal. A two-tailed test has a not
equal sign (≠) in the alternative hypothesis. When the population mean (µ) is not equal to
specified value of 𝜇0 , then alternative hypothesis can be expressed as µ ≠ 𝜇0 . A two-tailed
test is a hypothesis for which the rejection region lies on both ends of distribution, one on
the left and one on the right.
-For the test statistic, use the one-sample z test is used when there is only one sample in
the experiment that is known, and both the standard deviation and the mean of the
population are known also when the sample size is at least 30. On the other hand, to
compare a sample mean with the population mean when the population standard
deviation is not known, and when the sample size is less than 30, the one-sample t-test is
used.

Two-tailed Test One-tailed Test


Right-tailed Left-tailed
Sign in Ha ≠ > <
Rejection region Both sides Right side Left side

4. Determine the critical value and the rejection region/s.


- The critical values for a hypothesis test are the threshold to which the value of the test
statistic in a sample is compared to determine whether or not the null hypothesis is
rejected. The critical region or rejection, is a set of values of the test statistic for which the
null hypothesis is rejected in a hypothesis test, that is, the sample space for the test statistic
is partitioned into regions; the critical region will lead us to reject the null hypothesis, the
other not. Therefore, if the observed values of the test statistic is a member of the critical
region, we conclude “reject Ho”, if it is not a member of the critical region, then we can
conclude “do not reject Ho”.

GENERAL MATHEMATICS P a g e 4 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
- In here, we use the t-distribution table when we compute the critical value for a t-test
and the z-score when we compute the critical value of z. Recall your week 7 lesson in
midterm on how to get the t-tabular value and the z-score. (The t-distribution and z-score
are attached below)

5. Compute the test statistic.


-For t-test, use: -For z-test, use:

( 𝑥̄ − 𝝁)√𝒏 ( 𝑥̄ − 𝝁)√𝒏
𝒕𝒄𝒐𝒎𝒑𝒖𝒕𝒆𝒅 = 𝒛𝒄𝒐𝒎𝒑𝒖𝒕𝒆𝒅 =
𝒔 𝒔
Where, Where,
𝑥̄ = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑚𝑒𝑎𝑛 𝑥̄ = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑚𝑒𝑎𝑛
𝜇 = 𝑝𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑚𝑒𝑎𝑛 𝜇 = 𝑝𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑚𝑒𝑎𝑛
𝑛 = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑖𝑧𝑒 𝑛 = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑖𝑧𝑒
𝑠 = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑠 = 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛

6. State the decision rule.

GENERAL MATHEMATICS P a g e 5 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
- If the computed value of the statistic is greater than the critical value, then reject the null
hypothesis.
-If the computed value of the statistic is less than the tabular or critical value, then do not
reject the null hypothesis.

7. Conclusion
- Based on the test results, the researcher will reach a conclusion about the population
under study.

Let’s get started!

Since the steps in Hypothesis testing have been comprehensively discussed, you are now ready to
face the actual battle field as to how you are going to prove your claims. Before we proceed, always
take note that there are always two possible outcomes; if the result confirms the hypothesis, then
you’ve made measurements. If the result is contrary to the hypothesis, then you’ve made a
discovery. We may have not proved our claims, but at least, we discovered something. Who knows
if this discovery will have a great change and impact to our self, our belief and eventually, the
society we live in.

Example 1.
A test of the breaking strengths of six rings manufactured by a company showed a mean breaking
strength of 7750 pounds and a standard deviation of 145 pounds, whereas the manufacturer
claimed a mean breaking strength of 8000 pounds. Can we support the manufacturer’s claim at
significant level of 0.05?
Solution:

Ho:
a. State the null and µ = 8000 pounds and the manufacturer’s claim is justified
alternative hypotheses. Ha:
µ < 8000 pounds and the manufacturer’s claim is not justified
Since it has been stated from the problem that the significant level
b. Select the level of is 0.05 therefore,
significance
α = 0.05

c. Identify the appropriate since n=6 and 6 is less than 30 and the value of the population
test-statistic and mean is less than the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = t-test (one-tailed)

We use the t-distribution table since our test-statistic is t-test.

Given, α = 0.05

GENERAL MATHEMATICS P a g e 6 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
d. Determine the critical 𝑑𝑓 = 𝑛 − 1 = 6 − 1 = 5
value and the rejection
region/s. Refer on the table α = 0.05,
𝑑𝑓 = 5, one-tailed test

Critical value = -2.01

e. Compute the test Given: 𝑥̄ = 7750 𝑠 = 145 𝑛=6 𝜇 = 8000


statistic. Substitute.

( 𝑥̄ − 𝜇)√𝑛 ( 7750 − 8000)√6


𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = =
𝑠 145
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = −4.22

Since t-computed (|-4.22|) is greater than the t-tabular value (|-


2.01|), we shall reject the null hypothesis (Ho) and do not reject
f. State the decision rule.
alternative hypothesis (Ha).

*note that we always use the absolute value of the t-computed


and critical values when comparing them.

g. Conclusion The evidence is not enough to justify the manufacturer’s claim.

Example 2.
A researcher believes that in recent years, women have been getting taller. She knows that 10
years ago. The average height of young adult women living in her city was 63 inches. The standard
deviation is unknown. She randomly selected eight young adult women currently residing in her
city and measures their heights.
The following data are obtained: [64, 66, 68, 60, 62, 65, 66, and 63].

Solution:

a. State the null and Ho:


alternative hypotheses. µ = 63 the women in the city are not getting taller
Ha:
µ ≠ 63 inches and the women in the city is getting taller
b. Select the level of The significance level was not mentioned in the problem. In this
significance case, we use the 5%.

α = 0.05

c. Identify the appropriate since n=8 and 8 is less than 30 and the value of the population
test-statistic and mean is not equal to the specified value, therefore,

GENERAL MATHEMATICS P a g e 7 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
determine whether it is test statistic = t-test (two-tailed)
one-tailed or two-tailed.

d. Determine the critical We use the t-distribution table since our test-statistic is t-test.
value and the rejection
region/s. Given, α = 0.05

𝑑𝑓 = 𝑛 − 1 = 8 − 1 = 7

Refer on the table α = 0.05,


𝑑𝑓 = 5, two-tailed test

Critical value = ±2.365 (±, since it’s two tailed, the rejection
region is found at both ends of the normal curve)

e. Compute the test Solving for the mean and variance of the sample,
statistic. we have 𝑥̄ = 64.25 and 𝑠 2 = 6.5 then we get the square root
of the variance for us to solve the standard deviation.
Therefore, 𝑠 = 2.55
𝑥̄ = 64.25 𝑠 = 2.55 𝑛=8 𝜇 = 63

Substitute.
( 𝑥̄ − 𝜇)√𝑛 ( 64.25 − 63)√8
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = =
𝑠 2.55
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = 1.39
f. State the decision rule. Since t-computed (|1.39|) is less than the t-tabular value (|-
2.01|), we do not reject the null hypothesis (Ho) and reject
alternative hypothesis (Ha).

*note that we always use the absolute value of the t-computed


and critical values when comparing them

g. Conclusion There is no sufficient evidence that the women are getting taller
in the city.

Example 3.
According to the Department of Education, high school teachers work an average of 40 hours per
week during the school year. A district supervisor of a certain school surveyed 28 randomly
selected teachers and found out that they work an average of 42.6 hours a week and the standard
deviation was 3.75 hours. Test if the mean number of hours worked by teachers in the supervisor’s
school differs from the national average. Use α = 0.01
Solution:
a. State the null and Ho:
alternative hypotheses. µ = 40 hours and the hours worked by teachers in the supervisor’s
school does not differ from the national average

GENERAL MATHEMATICS P a g e 8 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Ha:
µ ≠ 40 hours and the hours worked by teachers in the supervisor’s
school differs from the national average
b. Select the level of Since it has been stated from the problem that the significant level
significance is 0.01 therefore,

α = 0.01

c. Identify the appropriate since n=28 and 28 is less than 30 and the value of the population
test-statistic and mean is not equal to the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = t-test (two-tailed)

d. Determine the critical We use the t-distribution table since our test-statistic is t-test.
value and the rejection
region/s. Given, α = 0.05

𝑑𝑓 = 𝑛 − 1
𝑑𝑓 = 28 − 1

𝑑𝑓 = 27

Refer on the table α = 0.05,


𝑑𝑓 = 27, two-tailed test

Critical value = ±2.771 (±, since it’s two tailed, the rejection region
is found at both ends of the normal curve)

e. Compute the test Given: 𝑥̄ = 42.6 𝑠 = 3.75 𝑛 = 28 𝝁 = 𝟒𝟎


statistic. Substitute.
( 𝑥̄ − 𝜇)√𝑛 ( 42.6 − 40)√28
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = =
𝑠 3.75
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = 3.67

f. State the decision rule. Since t-computed (|3.67|) is greater than the t-tabular value
(|±2.771 |), we reject the null hypothesis (Ho) and do not reject
alternative hypothesis (Ha).

*note that we always use the absolute value of the t-computed


and critical values when comparing them

g. Conclusion Therefore, there is sufficient evidence that the working hours of


28 teachers per week differ from the national average.

GENERAL MATHEMATICS P a g e 9 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Example 4.
A manufacturer claims that the average life of batteries used in their electronic games is 150 hours.
It is known that the standard deviation of this type of battery is 20 hours. A consumer wishes to
test the manufacturer’s claim and accordingly tests 100 electronic games using this battery and
found out that the mean is equal to 144 hours. Use α=0.05.
Solution:
a. State the null and Ho:
alternative hypotheses. µ = 150 hours (The average life of batteries used in electronic
games is equal to 150 hours)
Ha:
µ < 150 hours (The average life of batteries used in electronic
games less than 150 hours)
b. Select the level of Since it has been stated from the problem that the significant level
significance is 0.05 therefore,

α = 0.05

c. Identify the appropriate since n=100 and 100 is more than 30 and the value of the
test-statistic and population mean is less than the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = z-test (one-tailed)

d. Determine the critical We use the table for critical values of z since our test-statistic is z-
value and the rejection test.
region/s.
Given, α = 0.05

Refer on the table α = 0.05, one-tailed test

Critical value = ±1.645

e. Compute the test Given: 𝑥̄ = 144 ℎ𝑜𝑢𝑟𝑠 𝑠 = 20 𝑛 = 100 𝜇 = 150 ℎ𝑜𝑢𝑟𝑠


statistic. ( 𝑥̄ − 𝜇)√𝑛 ( 144 − 150)√100
𝑧𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = =
𝑠 20
𝑧𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = −3.00

f. State the decision rule. Since t-computed (|-3.00|) is greater than the z-critical value
(|±1.645 |), we reject the null hypothesis (Ho) and do not reject
alternative hypothesis (Ha).

*note that we always use the absolute value of the t-computed


and critical values when comparing them

g. Conclusion The evidence is not enough to prove that the life of batteries is
150 hours.

GENERAL MATHEMATICS P a g e 10 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Example 5
An achievement test was administered to thousands of pupils with mean score of 85 and a
standard deviation of 8. A random sample of 50 pupils were given the same test and showed an
average score of 83.20. Is there evidence to show that this group has lower performance than the
ones in general at 0.05 level of significance?
Solution:
a. State the null and Ho:
alternative hypotheses. µ = 85 (there is no evidence to show that the group has lower
performance than the ones in general)
Ha:
µ < 85 hours (the group has lower performance than the ones in
general)
b. Select the level of Since it has been stated from the problem that the significant level
significance is 0.05 therefore,

α = 0.05

c. Identify the appropriate since n=50 and 50 is more than 30 and the value of the population
test-statistic and mean is less than the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = z-test (one-tailed)

d. Determine the critical We use the table for critical values of z since our test-statistic is z-
value and the rejection test.
region/s.
Given, α = 0.05

Refer on the table α = 0.05, one-tailed test

Critical value = ±1.645

e. Compute the test Given: 𝑥̄ = 83.20 𝑠 = 8. 𝑛 = 50 𝜇 = 85


statistic. ( 83.2 − 85)√50
𝑧𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 =
8
𝑧𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = −1.59

f. State the decision rule. Since t-computed (|-1.59|) is less than the z-critical value (|±1.645
|), we do not reject the null hypothesis (Ho) and reject alternative
hypothesis (Ha).

*note that we always use the absolute value of the t-computed


and critical values when comparing them.

g. Conclusion There is no evidence to show that the group has lower


performance than the ones in general.

GENERAL MATHEMATICS P a g e 11 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Example 6.
A store owner claimed that the average weight of a pack of crackers is 250 g with a standard
deviation of 20 g. Would you agree to this claim if a random sample of 50 packs of crackers showed
an average of 242 g, using a 0.05 level of significance?
Solution:
a. State the null and Ho:
alternative hypotheses. µ = 250 (each pack of crackers do not weigh 250 g on the average)
Ha:
µ ≠ 250 (each pack of crackers weigh 250 g on the average)
b. Select the level of Since it has been stated from the problem that the significant level
significance is 0.05 therefore,

α = 0.05

c. Identify the appropriate Since n=50 and 50 is more than 30 and the value of the population
test-statistic and mean is not equal to the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = z-test (two-tailed)

d. Determine the critical We use the table for critical values of z since our test-statistic is z-
value and the rejection test.
region/s.
Given, α = 0.05

Refer on the table α = 0.05, two-tailed test

Critical value = ±1.96

e. Compute the test Given: 𝑥̄ = 242 𝑠 = 20. 𝑛 = 50 𝜇 = 250


statistic. ( 𝑥̄ − 𝜇)√𝑛 ( 242 − 250)√50
𝑧𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = =
𝑠 20
𝑧𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = −2.82

f. State the decision rule. Since t-computed (|-2.82|) is greater than the z-critical value
(|±1.96 |), we reject the null hypothesis (Ho) and do not reject
alternative hypothesis (Ha).

*note that we always use the absolute value of the t-computed


and critical values when comparing them.

g. Conclusion There is no evidence to show that each pack of crackers weigh 250 g.

GENERAL MATHEMATICS P a g e 12 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Generalizations.
Before we finally end this chapter, we summarize the ideas and the pieces of information with this
week’s lesson.
The z-test is a statistical hypothesis test that follows a normal distribution while T-test follows a
Student’s T-distribution.
The t-test is appropriate when you are handling small samples (n < 30) while z-test is appropriate
when you are handling moderate to large samples (n > 30). In decision making of both of the tests,
always remember that, if the computed value of the statistic is greater than the tabular or critical
value, reject the null hypothesis and accept the alternative hypothesis and if the computed value
of the statistic is less than the tabular or critical value, accept the null hypothesis and reject the
alternative hypothesis.
One thing that we can take away from this week’s lesson is that life is a hypothesis testing. You
don’t know the result but still you give it a shot because there’s a hope that you might land where
your results are. And in case it doesn’t work out, just the hypothesis of life.

GENERAL MATHEMATICS P a g e 13 | 13
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Finals – Second Semester Week 3: April 12-16, 2021

I. INTRODUCTION
This week, you shall be given another lesson to study and enrichment activities to solve.
Attached to this 8th week module is the weekly Study and Assessment Guide

DATE TOPIC ACTIVITIES OR TASKS


Hypothesis Testing of Two Means using Read and analyze the lessons.
April 12-16, 2021 Two-Sample z-test
Written Assessment 2

For this week of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the tasks
expected of you. GOOD LUCK!

Content Hypothesis Testing of Two Means using wo-Sample z-test


Learning Competencies At the end of the lesson, the learners would be able to:
• compare and contrast z-test and t-test;
• apply the formula of z-test in solving for the computed
value; and,
• test the difference between two means for independent
samples using the two-sample z-test.
Activities • Written Assessment 2
Essential Question/s • How can we accept and reject the decision that we
made?
• What are the consequences of your decisions?
Value Statement “Decisions are the hardest thing to make especially when it’s a
choice between where you should be and where you want to
be.” – M. Navarro
References Textbooks:
Rumsey. D (2021), Statistics Workbook For Dummies, Statistics II
For Dummies, and Probability For Dummies.

Myers, R., Walpole, R., [Link]. (2017). Probability & Statistics for
Engineers & Scientists, Pearson Education Limited, England.

Alferez, M. S. & Duro, M. C. A. (2018). Statistics and Probability.


MSA Publishing House. Cainta, Philippines.

De Guzman, D. (2017). Statistics and Probability. C & E


Publishing Inc. Quezon City, Philippines

Statistics and Probability P a g e 1 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
I. LEARNING CONTENT

Before we start our discussion let’s recall some of the basic terminologies that we need for today’s
discussion

Hypothesis Testing is a method that uses sample data to decide between two competing claims
about a population characteristic.

Hypothesis is a claim or statement either about the value of a single population characteristic or
about the values of several population.

The Null Hypothesis denoted by Ho is a claim about a population characteristic that is initially
assumed to be true.
The Alternative Hypothesis denoted by Ha is the competing claim.

Let’s recall the how to identify if what tailed-test we are going to apply using the given table.

Ho: parameter = specific value This is two-tailed test.


Ha: parameter ≠ specific value
Ho: parameter = specific value This is a left-tailed test (one-tailed).
Ha: parameter < specific value
Ho: parameter = specific value This is a right-tailed test (one-tailed).
Ha: parameter > specific value

Ho: 𝜇1 = 𝜇2 This is two-tailed test.


Ho: 𝜇1 ≠ 𝜇2
Ho: 𝜇1 ≤ 𝜇2
Ha: 𝜇1 > 𝜇2
Ho: 𝜇1 ≥ 𝜇2 This is one-tailed test
Ha: 𝜇1 < 𝜇2

Note: If the hypothesis is assuming that the two-sample means are equal, always use the two-
tailed test.
If the hypothesis is assuming that the one sample mean is greater than or less than the other
sample mean, always use the one-tailed test.

When do we use z-test and t-test?

The z-test and t-test are two of the statistical tools which can be used compare or to study
two groups of data through the value of their means.

What is the difference between z-test and t-test?

Statistics and Probability P a g e 2 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
The difference of t-test and z-test is depending on the number of sample size. If the
sample size is less than 30 thus the sample standard deviation is known, then we will use t-
test. However, we use
z-test when the number of sample size is greater than 30 or at least 30 and the population
standard deviation is known

Z test

When we use either one-tailed z-test or two -tailed z-test?

▪ One Sample Z-test is used when there is only one sample in the experiment that is
known, and both of the standard deviation and the mean of the population are known.
▪ Two Sample Z-Test is used when we want to compare the population means between
two independent sample groups.

For today’s discussion we will concentrate on Z-test specifically Two Sample Means.

Z test: Two Sample means formula.

̅
𝒙𝟏 − ̅
𝒙𝟐
𝒛=
(𝒔𝟏 )𝟐 (𝒔𝟐 )𝟐

𝒏𝟏 + 𝒏𝟐

𝑥̅1 → 𝑠𝑎𝑚𝑝𝑙𝑒 𝑚𝑒𝑎𝑛 𝑓𝑜𝑟 𝑔𝑟𝑜𝑢𝑝 1 𝑥̅2 → 𝑠𝑎𝑚𝑝𝑙𝑒 𝑚𝑒𝑎𝑛 𝑓𝑜𝑟 𝑔𝑟𝑜𝑢𝑝 1
𝑠1 → 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑓𝑜𝑟 𝑔𝑟𝑜𝑢𝑝 1 𝑠2 → 𝑠𝑎𝑚𝑝𝑙𝑒 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑓𝑜𝑟 𝑔𝑟𝑜𝑢𝑝 2
𝑛1 → 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑠𝑢𝑏𝑗𝑒𝑐𝑡𝑠 𝑖𝑛 𝑔𝑟𝑜𝑢𝑝 1 𝑛2 → 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑠𝑢𝑏𝑗𝑒𝑐𝑡𝑠 𝑖𝑛 𝑔𝑟𝑜𝑢𝑝 1

The Tabular Value of Z or the z-critical value at Indicated Level of Significance (𝜶)
Test / 𝜶 0.005 0.01 0.05 0.10
One-Tailed ±2.58 ±2.33 ±1.645 ±1.28
Two-Tailed ±2.81 ±2.575 ±1.96 ±1.645

General Procedures to be followed in using the z-test.

A. Formulate the null and the alternative hypothesis


B. Specify the level of significance and decide whether a one-tailed test or two-tailed test
C. Decide the test statistic to be used.
D. Compute for the value of the test statistics used using the sample data.
E. Make a decision

Statistics and Probability P a g e 3 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
➢ If the computed value of the test statistics is greater than the tabular or critical value,
reject the null hypothesis (Ho) and accept the alternative hypothesis (Ha) or do not
reject the alternative hypothesis.(Ha)
➢ If the computed value of the test statistics is less than the tabular or critical value,
accept the null hypothesis (Ho) or do not reject the null hypothesis (Ho) and reject the
alternative hypothesis. (Ha)
F. State the Conclusion.

For the purpose of discussion, we will use this problem:


Example #1
A bank is opening a branch in one of two neighborhoods, One of the factors considered
by the bank was whether the average monthly family income (in thousand pesos) in the two
neighborhoods differed. From census records, the bank drew two random samples of 100
families each and obtained the following information. a) The bank wishes to test the null
hypothesis that the two neighborhoods have the same mean income. b) The bank wishes to test
the hypothesis that neighbor A has a greater mean income than neighbor B.
What should the bank conclude? Use α = 0.05

Neighborhood
Sample A Sample B
𝑥̅1 = 10,100 𝑥̅2 = 10,300
𝑠1 = 300 𝑠2 = 400
𝑛1 = 100 𝑛2 = 100

Solution for a:
a) The bank wishes to test the null hypothesis that the two neighborhoods have the same
mean income.

Note: Since the hypothesis is assuming that the two neighbors have the same mean income,
we will use a two-tailed sample.
1st Step: Formulate the null and alternative 𝜇1 = 𝑇ℎ𝑒 𝑚𝑒𝑎𝑛 𝑖𝑛𝑐𝑜𝑚𝑒 𝑜𝑓 𝑛𝑒𝑖𝑔ℎ𝑏𝑜𝑟 𝐴
hypothesis.
𝜇2 = 𝑇ℎ𝑒 𝑚𝑒𝑎𝑛 𝑖𝑛𝑐𝑜𝑚𝑒 𝑜𝑓 𝑛𝑒𝑖𝑔ℎ𝑏𝑜𝑟 𝐵

Ho: 𝜇1 = 𝜇2 ; The two neighborhoods have


the same mean income.

Ha: 𝜇1 ≠ 𝜇2 ; The two neighborhoods do not


have the same mean income
nd
2 Step: Specify the Level of Significance and Level of Significance: 0.05
decide whether a one-tailed test or two- Since the problem requires to use two-tailed
tailed test. sample test, therefore the tabular value or
the critical value for 0.05 using two-tailed
sample tests is ±1.96

Statistics and Probability P a g e 4 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
3rd step: Decide the test statistic to be used. Z-test of Two Sample Means
4th Step: Compute for the value of the test statistics used using the sample data.

Identify the all of the given.


Neighborhood
Sample A Sample B
𝑥̅1 = 10,100 𝑥̅2 = 10,300
𝑠1 = 300 𝑠2 = 400
𝑛1 = 100 𝑛2 = 100
Use the formula: 𝑥̅1 − 𝑥̅2
𝑧=
( 𝑠 )2 ( 𝑠 ) 2
√ 1 + 2
𝑛 𝑛 1 2
Substitute the given on the formula and 𝑥̅1 − 𝑥̅2
𝑧=
simplify (𝑠1 )2 (𝑠2)2

𝑛1 + 𝑛2
10,100 − 10,300
𝑧= → 𝑆𝑖𝑚𝑝𝑙𝑖𝑓𝑦
2 2
√(300) + (400)
100 100
−200
𝑧=
√900 + 1600
−200
𝑧=
50
𝑧 = −4 → 𝐶𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒
5th Step: Make a decision. To make a decision we need to compare the
absolute value of computed value and the
absolute value of tabular value or the z-
critical value of the specified level of
significance used in the problem.

|𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒| ____ | 𝑧 −


𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 𝑣𝑎𝑙𝑢𝑒|
| − 4| ____ | ± 1.96|
| − 4| > | ± 1.96|

Since the computed value (|-4|) is greater


than the z-critical value (|±1.96|), we reject
the null hypothesis (Ho).

6th Step: State the conclusion The evidence is insufficient to prove that the
two neighborhoods have the same mean
income.

Solution for b:

Statistics and Probability P a g e 5 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
b) The bank wishes to test the hypothesis that neighbor A has a greater mean income than
neighbor B.

Note: Since the hypothesis is assuming that the neighbor A has a greater mean income that
neighbor B we will use one-tailed test.

Thus, Our Ha is always our claim.


1st Step: Formulate the null and alternative 𝜇1 = 𝑇ℎ𝑒 𝑚𝑒𝑎𝑛 𝑖𝑛𝑐𝑜𝑚𝑒 𝑜𝑓 𝑛𝑒𝑖𝑔ℎ𝑏𝑜𝑟 𝐴
hypothesis.
𝜇2 = 𝑇ℎ𝑒 𝑚𝑒𝑎𝑛 𝑖𝑛𝑐𝑜𝑚𝑒 𝑜𝑓 𝑛𝑒𝑖𝑔ℎ𝑏𝑜𝑟 𝐵

Ho: 𝜇1 ≤ 𝜇2 ; The mean income of neighbor A


is less than or equal to the mean income of
neighbor B.

Ha: 𝜇1 > 𝜇2 (claim): The mean income of


neighbor A is greater than the mean income
of neighbor B.
nd
2 Step: Specify the Level of Significance and Level of Significance: 0.05
decide whether a one-tailed test or two- Since the problem requires to use one-tailed
tailed test. sample test, therefore the tabular value or
the z-critical value for 0.05 using one-tailed
sample tests is ±1.645
3rd step: Decide the test statistic to be used. Z-test of Two Sample Means
th
4 Step: Compute for the value of the test statistics used using the sample data.

Identify the all of the given. Neighborhood


Sample A Sample B
𝑥̅1 = 10,100 𝑥̅2 = 10,300
𝑠1 = 300 𝑠2 = 400
𝑛1 = 100 𝑛2 = 100
Use the formula: 𝑥̅1 − 𝑥̅2
𝑧=
( 𝑠 )2 ( 𝑠 ) 2
√ 1 + 2
𝑛 𝑛 1 2
Substitute the given on the formula and 𝑥̅1 − 𝑥̅2
𝑧=
simplify (𝑠1 )2 (𝑠2)2

𝑛1 + 𝑛2
10,100 − 10,300
𝑧= → 𝑆𝑖𝑚𝑝𝑙𝑖𝑓𝑦
2 2
√(300) + (400)
100 100
−200
𝑧=
√900 + 1600
−200
𝑧=
50

Statistics and Probability P a g e 6 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
𝑧 = −4 → 𝐶𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒
5th Step: Make a decision. To make a decision we need to compare the
absolute value of computed value and the
absolute value of tabular value or the z-
critical value of the specified level of
significance used in the problem.

|𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒| ____ | 𝑧 −


𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 𝑣𝑎𝑙𝑢𝑒|
| − 4| ____ | ± 1.645|
| − 4| > | ± 1.645|

Since the computed (|-4|) is greater than the


z-critical value (|±1.645|), we reject the null
hypothesis (Ho).

6th Step: State the conclusion The two neighborhoods do not have the
same mean income.

Example #2

Mr. Joe Rivera, a senior high school teacher claims that the students in his class will get a higher
score on the Statistics Examination than Ms. Marie Santiago’s class. The mean of Mr. Joe’s class
for 51 students is 23.3 and the standard deviation is 4.7 while the mean of Ms. Marie’s class for
47 students is 18.3 and the standard deviation of 5.3. Using α = 0.10 can the claim of Mr. Joe
be true?

1st Step: Formulate the null and alternative 𝜇1 = 𝑇ℎ𝑒 𝑠𝑐𝑜𝑟𝑒 𝑜𝑓 𝑀𝑟. 𝐽𝑜𝑒 ′ 𝑠 𝑐𝑙𝑎𝑠𝑠
hypothesis.
𝜇2 = 𝑇ℎ𝑒 𝑠𝑐𝑜𝑟𝑒 𝑜𝑓 𝑀𝑠. 𝑀𝑎𝑟𝑖𝑒 ′ 𝑠 𝑐𝑙𝑎𝑠𝑠

Ho: 𝜇1 ≤ 𝜇2 ; The scores of Mr. Joe’s class is


lower than or equal to the scores of Ms.
Marie’s class.

Ha: 𝜇1 > 𝜇2 (claim): The scores of Mr. Joe’s


class is higher than the scores of Ms. Marie’s
class. (claim)

2nd Step: Specify the Level of Significance and Level of Significance: 0.10
decide whether a one-tailed test or two- Since the problem requires to use one-tailed
tailed test. sample test, therefore the tabular value or
the z-critical value for 0.10 using one-tailed
sample tests is ±1.28
3rd step: Decide the test statistic to be used. Z-test of Two Sample Means

Statistics and Probability P a g e 7 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
4th Step: Compute for the value of the test statistics used using the sample data.

Identify the all of the given. Mr. Joe’s Class Ms. Marie’s Class
𝑥̅1 = 23.3 𝑥̅2 = 18.3
𝑠1 = 4.7 𝑠2 = 5.3
𝑛1 = 51 𝑛2 = 47
Use the formula: 𝑥̅1 − 𝑥̅2
𝑧=
(𝑠1 )2 (𝑠2 )2

𝑛1 + 𝑛2
Substitute the given on the formula and 𝑥̅1 − 𝑥̅2
𝑧=
simplify (𝑠1 )2 (𝑠2)2

𝑛1 + 𝑛2
23.3 − 18.3
𝑧= → 𝑆𝑖𝑚𝑝𝑙𝑖𝑓𝑦
2 2
√(4.7) + (5.3)
51 47
5
𝑧=
√0.433 + 0.598
1
𝑧=
1.015
𝑧 = 0.985 → 𝐶𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒
5th Step: Make a decision. To make a decision we need to compare the
absolute value of computed value and the
absolute value of tabular value or the z-
critical value of the specified level of
significance used in the problem.

|𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒| ____ | 𝑧 −


𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 𝑣𝑎𝑙𝑢𝑒|
|0.985| ____ | ± 1.28|
|0.985| < | ± 1.28|

Since the computed (|0.985|) is less than the


z-critical value (|±1.28|), do not reject the
null hypothesis (Ho).
6th Step: State the conclusion The scores of Mr. Joe’s class is lower than or
equal to the scores of Ms. Marie’s class.

Example #3
The average weight of 50 Volleyball Players who are properly trained for 3 to 6 months was 76.8
kilogram with the standard deviation of 12.6 kilogram while the other 38 volleyball players who
were trained only for less than 3 months have an average weight of 72.5 kilogram with the

Statistics and Probability P a g e 8 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
standard deviation of 10.7 kilogram. Test the hypothesis that the Volleyball Players who are
properly trained have an equal weight to the volleyball players who were trained for less than 3
months. Use α = 0.05 level of Significance.

1st Step: Formulate the null and alternative 𝜇1 = The weight of volleyball players who
hypothesis. were properly trained for 3 to 6 months.

𝜇2 =The weight of the volleyball players who


were trained for less than 3 months.

Ho: 𝜇1 = 𝜇2 ; The weight of volleyball players


who were properly trained for 3 to 6 months
have an equal weight to the volleyball players
who were trained for less than 6 months .
(claim)

Ha: 𝜇1 ≠ 𝜇2 : The weight of volleyball players


who were properly trained for 3 to 6 months
is not equal weight to the volleyball players
who were trained for less than 6 months

2nd Step: Specify the Level of Significance and Level of Significance: 0.05
decide whether a one-tailed test or two- Since the problem requires to use two-tailed
tailed test. sample test, therefore the tabular value or
the z-critical value for 0.05 using two-tailed
sample tests is ±1.96
3rd step: Decide the test statistic to be used. Z-test of Two Sample Means
th
4 Step: Compute for the value of the test statistics used using the sample data.

Identify the all of the given. VB Players who VB Players who


were Trained for 3-6 were Trained for
months less than 3 months
𝑥̅1 = 76.8 𝑥̅2 = 72.5
𝑠1 = 12.6 𝑠2 = 10.7
𝑛1 = 50 𝑛2 = 38
Use the formula: 𝑥̅1 − 𝑥̅2
𝑧=
( 𝑠 )2 ( 𝑠 ) 2
√ 1 + 2
𝑛 𝑛 1 2
Substitute the given on the formula and 𝑥̅1 − 𝑥̅2
𝑧=
simplify (𝑠1 )2 (𝑠2)2

𝑛1 + 𝑛2
76.8 − 72.5
𝑧= → 𝑆𝑖𝑚𝑝𝑙𝑖𝑓𝑦
( )2 ( )2
√ 12.6 + 10.7
50 38

Statistics and Probability P a g e 9 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
2.3
𝑧=
√3.175 + 3.013
4.3
𝑧=
2.488
𝑧 = 1.728 → 𝐶𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒
5th Step: Make a decision. To make a decision we need to compare the
absolute value of computed value and the
absolute value of tabular value or the z-
critical value of the specified level of
significance used in the problem.

|𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒| ____ | 𝑧 −


𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 𝑣𝑎𝑙𝑢𝑒|
|1.728| ____ | ± 1.96|
|1.728| < | ± 1.96|

Since the computed (|1.782|) is less than the


z-critical value (|±1.96|), do not reject the
null hypothesis (Ho).
6th Step: State the conclusion The weight of volleyball players who were
properly trained for 3 to 6 months have an
equal weight to the volleyball players who
were trained for less than 6 months.

Example #4
The IQs of 32 students from the section A showed a mean of 107 with a standard deviation of
10. While the IQs of 41 students from the section B showed a mean of 112 with a standard
deviation of 8. Is there a significant difference between the IQs of St. Anselm and St. Augustine?
Use 0.01 level of significance.

1st Step: Formulate the null and alternative 𝜇1 = The IQ of section A


hypothesis.
𝜇2 = The IQ of section B

Ho: There is no significant difference between


the IQs of St. Anselm and St. Augustine

Ha: There is a significant difference between


the IQs of the two sections. (claim)

2nd Step: Specify the Level of Significance and Level of Significance: 0.01
decide whether a one-tailed test or two- Since the problem requires to use two-tailed
tailed test. sample test, therefore the tabular value or

Statistics and Probability P a g e 10 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
the z-critical value for 0.01 using two-tailed
sample tests is ±2.575
rd
3 step: Decide the test statistic to be used. Z-test of Two Sample Means
th
4 Step: Compute for the value of the test statistics used using the sample data.

Identify the all of the given. Section A Section A


𝑥̅1 = 107 𝑥̅2 = 112
𝑠1 = 10 𝑠2 = 8
𝑛1 = 32 𝑛2 = 41
Use the formula: 𝑥̅1 − 𝑥̅2
𝑧=
(𝑠1 )2 (𝑠2 )2

𝑛1 + 𝑛2
Substitute the given on the formula and 𝑥̅1 − 𝑥̅2
𝑧=
simplify (𝑠1 )2 (𝑠2)2

𝑛1 + 𝑛2
107 − 112
𝑧= → 𝑆𝑖𝑚𝑝𝑙𝑖𝑓𝑦
2 2
√(10) + (8)
32 41
−5
𝑧=
√3.125 + 1.561
−5
𝑧=
2.165
𝑧 = −2.310 → 𝐶𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒
5th Step: Make a decision. To make a decision we need to compare the
absolute value of computed value and the
absolute value of tabular value or the z-
critical value of the specified level of
significance used in the problem.

|𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒| ____ | 𝑧 −


𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 𝑣𝑎𝑙𝑢𝑒|
| − 2.310| ____ | ± 2.575|
|−2.310| < | ±2.575|

Since the computed (|-2.310|) is less than


the z-critical value (|±2.575|), do not reject
the null hypothesis (Ho).
6th Step: State the conclusion There is no significant difference between the
IQs of the two sections.

Statistics and Probability P a g e 11 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Generalization:

• If the number of samples is greater than 30 or at least 30 and the population standard
deviation is known, use Z test.
• Two Sample Z-Test is used when we want to compare the population means between
two groups.
• |𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒| < | 𝑧 − 𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 𝑣𝑎𝑙𝑢𝑒|, do not reject the null hypothesis (Ho).
• |𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒| > | 𝑧 − 𝑐𝑟𝑖𝑡𝑖𝑐𝑎𝑙 𝑣𝑎𝑙𝑢𝑒|, reject the null hypothesis (Ho).

Statistics and Probability P a g e 12 | 12


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Final Term – Second Semester Week 4: April 19-23, 2021

I. INTRODUCTION

This week, you shall be given another lesson to study and enrichment activities to solve.
Attached to this 4th week module is the weekly Study and Assessment Guide.

DATE TOPIC ACTIVITIES OR TASKS


Hypothesis Testing Two Sample t-Test Read and analyze the lessons
April 19-23, 2021
Written Assessment 3

For these weeks of the final term, the following shall be your guide for the different
lessons and tasks that you need to accomplish. Be patient, read it carefully before proceeding to
the tasks expected of you.
GOOD LUCK!

Content Hypothesis Testing for Two Sample T-test


Learning Competencies At the end of the lesson, the learners would be able to:
• compare and contrast z-test and t-test;
• apply the formula of t-test in solving for the computed
value; and,
• test the difference between two means for independent
samples using the two-sample t-test.
Activities • Written Assessment 3
Essential Questions • Why do we need to test hypothesis?
Value Statement • Decisions must be based on empirical evidences not simply
on opinion or hearsay.
References Textbooks:
Rumsey. D (2021), Statistics Workbook For Dummies, Statistics II
For Dummies, and Probability For Dummies.

Myers, R., Walpole, R., [Link]. (2017). Probability & Statistics for
Engineers & Scientists, Pearson Education Limited, England.

Alferez, M. S. & Duro, M. C. A. (2018). Statistics and Probability.


MSA Publishing House. Cainta, Philippines.

De Guzman, D. (2017). Statistics and Probability. C & E Publishing


Inc. Quezon City, Philippines

Statistics and Probability P a g e 1|8


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
II. LEARNING CONTENT

In the previous week, the z test was used to test the difference between two means when the
population standard deviation was known and the variables were normally or approximately
normally distributed, or when both sample sizes were greater than or equal to 30. In many
situations, however, these conditions cannot be met-that is, the population standard deviations
are not known. In these cases, a t-test is used to test the difference between means when the
two sample are dependent and when the samples are taken from two normally or approximately
normally distributed populations. Samples are independent samples when they are not related.

Formula for the t-test for Testing the Difference Between Two Means-Independent Samples
Variances are assumed to be unequal:
(𝑥1 − 𝑥2 ) − (𝜇1 − 𝜇2 )
𝑡=
(𝑠12 ) (𝑠22 )
√ +
𝑛1 𝑛2

where:
𝑠1and 𝑠2, the sample standard deviations, are estimates of 𝜎1 and 𝜎2 , respectively.
𝑥1 and 𝑥2 are the sample means.
𝜇1 and 𝜇2 are the population means.

The degrees of freedom are equal to the smaller of 𝑛1 + 𝑛2 − 2.

The formula
(𝒙𝟏 − 𝒙𝟐 ) − (𝝁𝟏 − 𝝁𝟐 )
𝒕=
(𝒔𝟐𝟏 ) (𝒔𝟐𝟐 )

𝒏𝟏 + 𝒏𝟐

follows the format of


𝑜𝑏𝑠𝑒𝑟𝑣𝑒𝑑 𝑣𝑎𝑙𝑢𝑒 − 𝑒𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝑣𝑎𝑙𝑢𝑒
𝑇𝑒𝑠𝑡 𝑉𝑎𝑙𝑢𝑒 =
𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑒𝑟𝑟𝑜𝑟

Where 𝑥1 − 𝑥2 is the observed difference between sample means and where the expected value
𝜇1 − 𝜇2 is equal to zero when no difference between population mean is hypothesized. The
(𝑠2 ) (𝑠22)
denominator √ 𝑛1 + is the standard error of the difference between two means. Since
1 𝑛2
mathematical deviation of the standard error is somewhat complicated, it will be omitted here.

Example1:
The average size of a farm in Indiana Country, Pennsylvania is 191 acres. The average size of a
fam in Greene Country, Pennsylvania is 199 acres. Assume the data were obtained from two

Statistics and Probability P a g e 2|8


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
samples with standard deviations of 38 and 12 acres, respectively, and sample sizes of 8 and 10,
respectively. Can it be concluded at 𝛼 = 0.05 that the average size of the farm in the two
countries is different? Assume the populations are normally distributed.

Solution:
a. State the null and alternative hypothesis.
𝐻0: 𝜇1 = 𝜇2
𝐻𝑎: 𝜇1 ≠ 𝜇2

b. Select the level of significance.


𝛼 = 0.05

c. Identify the test-statistic.


T-test for two Independent Sample Means

d. Determine the critical value and the rejection region/s.


Since the test is two-tailed, since 𝛼 = 0.05, and since the variances are unequal, the
degrees of freedom are the smaller number 𝑛1 + 𝑛2 − 2. In this case, the degrees of
freedom are 10 + 8 − 2 = 16. Hence from the t-distribution table, the critical values are
+2.120 and −2.120.
(See t-distribution table attached to week7 module during the midterms)

Confidence 0.500 0.800 0.900 0.950 0.980 0.990


interval
d.f. One tail 𝒕.𝟐𝟓𝟎 𝒕.𝟏𝟎𝟎 𝒕.𝟎𝟓𝟎 𝒕.𝟎𝟐𝟓 𝒕.𝟎𝟏𝟎 𝒕.𝟎𝟎𝟓
d.f. Two-tail 𝒕.𝟓𝟎𝟎 𝒕.𝟐𝟎𝟎 𝒕.𝟏𝟎𝟎 𝒕.𝟎𝟓𝟎 𝒕.𝟎𝟐𝟎 𝒕.𝟎𝟏𝟎
1
2
3

16 2.120

e. Compute the test statistic.


Since the variance are unequal, we use the formula,
Thus,
(𝑥1 − 𝑥2 ) − (𝜇1 − 𝜇2 ) (191 − 199) − 0 −8 −8
𝑡= = = = = 𝟎. 𝟓𝟕
2 2 (38 2) (122) 1444 144 13.96
(𝑠 ) (𝑠 ) √ √
√ 1 + 2 8 + 10 8 + 10
𝑛1 𝑛2

Statistics and Probability P a g e 3|8


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
f. State the decision rule.
If t is less than -2.120 or greater than 2.120, reject the null hypothesis.
Since 𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 (−0.57) > 𝑡∝ (−2.120), do not reject 𝐻𝑜 .
2
g. State the Conclusion
There is no enough evidence to support the claim that the average size of the farm is
different.

Example 2: A Statistics teacher wants to compare his two classes to see if the performed any
differently on Midterm examination. Class A had 25 students with an average score of 70,
standard deviation of 15. Class B had 20 students with an average score of 74, standard deviation
of 25. Using alpha 0.05, did theses two classes perform differently on the test?

Solution:
a. State the null and alternative hypothesis.
𝐻0: 𝜇𝐴 = 𝜇𝐵
𝐻𝑎: 𝜇𝐴 ≠ 𝜇𝐵

b. Select the level of significance.


∝= 0.05

c. Identify the test-statistic.


T-test for two Independent Sample Means

d. Determine the critical value and the rejection region/s.


Since the test is two-tailed, since 𝛼 = 0.05, and since the variances are unequal, the
degrees of freedom are the smaller number 𝑛1 + 𝑛2 − 1. In this case, the degrees of
freedom are 25 + 20 − 2 = 43. Hence from the t-distribution table, the critical values
are +1.960 and −1.960.

e. Compute the test statistic.


Since the variance are unequal, we use the formula,

(𝑥1 − 𝑥2 ) − (𝜇1 − 𝜇2 ) (70 − 74) − (0) −4


𝑡= = = = −0.6305
2 2
(𝑠 2 ) (𝑠22 ) √(15) + (25) √161
√ 1 25 20 4
𝑛1 + 𝑛2

f. State the decision rule.


If t is less than -1.960 or greater than 1.960, reject the null hypothesis.
Since 𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 (−0.6305) > (−1.960), do not reject 𝐻𝑜 .
g. State the Conclusion
There is no significant difference on the test performance of Class A and Class B.

Statistics and Probability P a g e 4|8


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Example 3: Mr. Henrick grows mushroom using two different type of soil. When the mushroom
is ready for harvest, he is curious as to whether the sizes of his tomato plants differ with the two
type of soil used. He takes a random sample from each type of soil and weighs the mushroom.
Here is the summary of the result.
TYPE 1 TYPE 2
Mean 1.50 m 1.36 m
Standard Deviation 0.67 m 0. 75 m
Number of Mushroom 23 25

At 𝑎𝑙𝑝ℎ𝑎 = 0.05, can we conclude that the sizes of the mushrooms differ between the type of
soils used?

Solution:
a. State the null and alternative hypothesis.
𝐻0: 𝜇𝑡𝑦𝑝𝑒1 = 𝜇𝑡𝑦𝑝𝑒𝑏
𝐻𝑎: 𝜇𝑡𝑝𝑒1 ≠ 𝜇𝑡𝑦𝑝𝑒2

b. Select the level of significance.


𝛼 = 0.05

c. Identify the test-statistic.


T-test for two Independent Sample Means

d. Determine the critical value and the rejection region/s.


Since the test is two-tailed, since 𝛼 = 0.05, and since the variances are unequal, the
degrees of freedom are the smaller number 𝑛1 + 𝑛2 − 2. In this case, the degrees of
freedom are 23 + 25 − 2 = 46. Hence from the t-distribution table, the critical values
are +1.960 and −1.960.

e. Compute the test statistic.


Since the variance are unequal, we use the formula,

(𝑥1 − 𝑥2 ) − (𝜇1 − 𝜇2 ) (1.50 − 1.36) − (0) −4


𝑡= = = = 𝟑. 𝟑𝟑𝟐
( )2 ( )2 √161
(𝑠 2 ) (𝑠22 ) √ 0.67 + 0.75
√ 1 23 25 4
𝑛1 + 𝑛2
f. State the decision rule.
If t is less than -1.960 or greater than 1.960, reject the null hypothesis.
Since 𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 (3.332) > (1.960), reject 𝐻𝑜 .
g. State the Conclusion
There is a significant difference on the sizes of mushroom when planted on two different
type of soil.

Statistics and Probability P a g e 5|8


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Example 4: The average time boys and girls aged three to eight spend studying each day is
believed to be the same. A study is done and data are collected, resulting in the data in the table
below.

Sample Average Number of Hours Studying Sample Standard Deviation

Girls 9 2 0.866

Boys 16 3.2 1.00

Assume that the data is normally distributed, Is there a difference in the mean amount of time
boys and girls aged three to eight study each day? Test at the 5% level of significance.

Solution:

a. State the null and alternative hypothesis.


𝐻0: 𝜇𝑔𝑖𝑟𝑙𝑠 = 𝜇𝑏𝑜𝑦𝑠
𝐻𝑎: 𝜇𝑔𝑖𝑟𝑙𝑠 ≠ 𝜇𝑏𝑜𝑦𝑠

b. Select the level of significance.


𝛼 = 0.05

c. Identify the test-statistic.


T-test for two Independent Sample Means

d. Determine the critical value and the rejection region/s.


Since the test is two-tailed, since ∝= 0.05, and since the variances are unequal, the
degrees of freedom are the smaller number 𝑛1 + 𝑛2 − 2. In this case, the degrees of
freedom are 9 + 16 − 2 = 23. Hence from the t-distribution table, the critical values are
+2.069 and −2.069.

e. Compute the test statistic.


Since the variance are unequal, we use the formula,

(𝑥1 − 𝑥2 ) − (𝜇1 − 𝜇2 ) (2 − 3.2) − (0) −1.2


𝑡= = = = −3.142
2 2
(𝑠 2 ) (𝑠22 ) √(0.866) + (1.00) √ 164057
√ 1 9 16 1125000
𝑛1 + 𝑛2

f. State the decision rule.


If t is less than -2.069 or greater than 2.069, reject the null hypothesis.
Since 𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 (−3.142) < (−2.069), Reject 𝐻𝑜 .
g. State the Conclusion

Statistics and Probability P a g e 6|8


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
At 5% level of significance, the sample data shows that there is sufficient evidence to
conclude that the mean number of hours that girls and boys aged three to eight spent in
studying is significantly different.

Example 5: A US Magazine, Consumer Reports, carried out a survey of the calorie content of a
number of different brand of hotdogs. The calorie content of 20 beef and 17 poultry hotdogs was
recorded as below:

Average Calorie Content Sample Standard Deviation

Beef 156.9 22.6

Poultry 122.5 25.5

Is there a difference in calorie content between beef and poultry hotdogs? Use 5% level of
significance.

Solution:

a. State the null and alternative hypothesis.


𝐻0: 𝜇𝑏𝑒𝑒𝑓 = 𝜇𝑝𝑜𝑟𝑘
𝐻𝑎: 𝜇𝑏𝑒𝑒𝑓 ≠ 𝜇𝑝𝑜𝑟𝑘

b. Select the level of significance.


𝛼 = 0.05

c. Identify the test-statistic.


T-test for two Independent Sample Means

d. Determine the critical value and the rejection region/s.


Since the test is two-tailed, since 𝛼 = 0.05, and since the variances are unequal, the
degrees of freedom are the smaller number 𝑛1 + 𝑛2 − 2. In this case, the degrees of
freedom are 20 + 17 − 2 = 35. Hence from the t-distribution table, the critical values
are +1.960 and −1.960.

e. Compute the test statistic.


Since the variance are unequal, we use the formula,

(𝑥1 − 𝑥2 ) − (𝜇1 − 𝜇2 ) (156.9 − 122.5) − (0) 34.4


𝑡= = = = 4.307
2 2
(𝑠 2 ) (𝑠22 ) √(22.6) + (25.5) √15947
√ 1 20 17 250
𝑛1 + 𝑛2
f. State the decision rule.
If t is less than -1.960 or greater than 1.960, reject the null hypothesis.
Since 𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 (4.307) > (1.960), Reject 𝐻𝑜 .

Statistics and Probability P a g e 7|8


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
g. State the Conclusion
There is very strong evidence that the calorie content of hot dogs for these two groups is
different

Generalization
The two-sample t-test (also known as the independent samples t-test) is a method used to test
whether the unknown population means of two groups are equal or not. It is used if the following
conditions are met.
1. Data values are independent.
2. Data values are randomly sampled from two normal populations.
3. The two individual groups have equal variances.

Statistics and Probability P a g e 8|8


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Final Term – Second Semester Week 5: April 26-30, 2021

I. INTRODUCTION

Hello, Louisian GEM! I hope you are doing great as we unravel the facts and some pieces
of important information of this week’s corresponding learning. Also, keep the spirit and continue
soaring high for the remaining weeks of this school year.

This week, you shall be given another lesson to study and learning task to accomplish.
Attached to this module is the weekly Study and Assessment Guide

DATE TOPIC ACTIVITIES OR TASKS


Hypothesis Testing: Read the lessons.
April 26-30, 2021
Paired Sample t-test Answer the learning tasks

For this week of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the tasks
expected of you. GOOD LUCK!

Content Hypothesis Testing:


Paired Sample t-test
Learning Competencies The learner should be able to:
• illustrate the steps in hypothesis testing using paired
sample t-test;
• apply the formula in getting the computed test statistic
using the paired sample t-test;
• identify the parameter to be tested given a real-life
problem; and,
• draw conclusion about the population mean based on the
test-statistic value and the rejection region.
Activities Written Assessment 4
Essential Questions Why do you think it is important to determine if you have improved
on two sets of tests?
Value Statement “No matter how good you get you can always get better, and that's
the exciting part.”
- Tiger Woods
References Textbooks:
Jimenez, R. et, al. (2014). Basic Statistics a Worktext. C & E
Publishing, [Link] City, Philippines

Rumsey. D (2021), Statistics Workbook for Dummies, Statistics II For


Dummies, and Probability For Dummies.

PROBABILITY AND STATISTICS P a g e 1 | 11


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Myers, R., Walpole, R., [Link]. (2017). Probability & Statistics for
Engineers & Scientists, Pearson Education Limited, England.

Alferez, M. S. & Duro, M. C. A. (2018). Statistics and Probability. MSA


Publishing House. Cainta, Philippines.

De Guzman, D. (2017). Statistics and Probability. C & E Publishing


Inc. Quezon City, Philippines
Bluman, A. (2009). Elementary Statistics A Step-by-Step Approach.
The McGraw-Hill Companies, Inc., New York City, USA

II. LEARNING CONTENT

On the previous weeks, you have learned about hypothesis testing and some statistical tools to
be used. In this section, a different version of the t- test is explained. This version is used when the
samples are dependent. Samples are considered to be dependent samples when the subjects are
paired or matched in some way.

Did you observe that your teacher, sometimes conducts a pre-test then a post-test? Why do you
think he did that? Simply because he wants to determine if you have an improvement or your
performance on your pre-test differs from your post- test.

What do you think your teacher do to determine if you have improved on that certain two sets
of tests? Well, the appropriate test is a paired sample t-test.

Now, let us define a paired sample t-test.

What is a paired sample t-test?

The paired sample t-test, sometimes called the dependent sample t-test, is a statistical procedure
used to determine whether the mean difference between two sets of observations is zero. In a
paired sample t-test, each subject or entity is measured twice, resulting in pairs of observations.
Common applications of the paired sample t-test include case-control studies or repeated-
measures designs. Suppose you are interested in evaluating the effectiveness of a company
training program. One approach you might consider would be to measure the performance of a
sample of employees before and after completing the program, and analyze the differences using
a paired sample t-test.

For example, suppose a medical researcher wants to see whether a drug will affect the reaction
time of its users. To test this hypothesis, the researcher must pretest the subjects in the sample
first. That is, they are given a test to ascertain their normal reaction times. Then after taking the
drug, the subjects are tested again, using a post-test. Finally, the means of the two tests are
compared to see whether there is a difference. Since the same subjects are used in both cases the

PROBABILITY AND STATISTICS P a g e 2 | 11


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
samples are related; subjects scoring high on the pretest will generally score high on the posttest,
even after consuming the drug. Likewise, those scoring lower on the pretest will tend to score
lower on the posttest. To take this effect into account, the researcher employs a t- test, using the
differences between the pretest values and the posttest values. Thus, only the gain or loss in values
is compared.

How do we formulate the hypotheses in paired sample t-test?

The paired sample t-test hypotheses are formally defined below:

• The null hypothesis (Ho) assumes that the true mean difference (μd) is equal to zero.
• The two-tailed alternative hypothesis (Ha) assumes that μd is not equal to zero.
• The upper-tailed alternative hypothesis (Ha) assumes that μd is greater than zero.
• The lower-tailed alternative hypothesis (Ha) assumes that μd is less than zero.

The mathematical representations of the null and alternative hypotheses are defined below:

Two- tailed Left- tailed Right- tailed


𝐻𝑜 : 𝜇𝑑 = 0 𝐻𝑜 : 𝜇𝑑 = 0 𝐻𝑜 : 𝜇𝑑 = 0
𝐻𝑜 : 𝜇𝑑 ≠ 0 𝐻𝑜 : 𝜇𝑑 < 0 𝐻𝑜 : 𝜇𝑑 > 0

Note. It is important to remember that hypotheses are never about data, they are about the
processes which produce the data. In the formulas above, the value of μd is unknown. The goal of
hypothesis testing is to determine the hypothesis (null or alternative) with which the data are
more consistent.

Before we proceed to answering some of the sample problems, we recall first the steps in
Hypothesis testing which have been discussed during the 1st week of your module.

Review of the Steps in Hypothesis Testing

1. State the null and alternative hypotheses.


The mathematical representations of the null and alternative hypotheses are defined
below:

Two- tailed Left- tailed Right- tailed


𝐻𝑜 : 𝜇𝑑 = 0 𝐻𝑜 : 𝜇𝑑 = 0 𝐻𝑜 : 𝜇𝑑 = 0
𝐻𝑜 : 𝜇𝑑 ≠ 0 𝐻𝑜 : 𝜇𝑑 < 0 𝐻𝑜 : 𝜇𝑑 > 0

2. Select the level of significance.


-The significance level which is usually denoted by alpha (α) is related to the degree of
certainty we require in order to reject the null hypothesis in favor of the alternative
hypothesis. The most common significance levels are 0.05 and 0.01 because of the desire

PROBABILITY AND STATISTICS P a g e 3 | 11


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
to maintain a low probability of rejecting the null hypothesis when it is in fact true. The use
of a level of significance (α) of 0.05 or 0.01 means that we are willing to commit an error
of 5% or 1% and are, therefore, confident of making 95% or 99% correct decision. If,
instance, we set the (α)=0.5, the probability of incorrectly rejecting the null hypothesis
when it is in fact true at 5% can be avoided or protected from error by choosing a lower
value for α.
3. Identify the appropriate test-statistic and determine whether it is one-tailed or two-tailed.
-For the test statistic, use a paired sample t-test or t-test on the significance of the
difference between two correlated means. A paired samples t-test is used to compare the
means of two samples when each observation in one sample can be paired with an
observation in the other sample.

4. Determine the critical value and the rejection region/s.


In here, we use the t-distribution table when we determine the t critical values.
Recall how to get the t-tabular value. We still follow the decision rule as provided before.

PROBABILITY AND STATISTICS P a g e 4 | 11


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
5. Compute the test statistic.
-For paired sample t-test, use:
𝚺(𝒙𝟏 − 𝒙𝟐 )
− 𝝁𝒅
𝒕𝒄𝒐𝒎𝒑𝒖𝒕𝒆𝒅 = 𝒏
𝒏(𝚺𝒅𝟐) − (𝚺𝒅)𝟐

𝒏(𝒏 − 𝟏)
√𝒏
or
𝒅
− 𝝁𝒅
𝒕𝒄𝒐𝒎𝒑𝒖𝒕𝒆𝒅 = 𝒏 𝒔
𝒅
√𝒏
where,
𝑑 = 𝑑𝑖𝑓𝑓𝑒𝑟𝑒𝑛𝑐𝑒 𝑏𝑒𝑡𝑤𝑒𝑒𝑛 𝑡ℎ𝑒 𝑣𝑎𝑙𝑢𝑒𝑠 𝑜𝑓 𝑡ℎ𝑒 𝑝𝑎𝑖𝑟𝑠 𝑜𝑓 𝑑𝑎𝑡𝑎
𝑛 = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑜𝑏𝑠𝑒𝑟𝑣𝑎𝑡𝑖𝑜𝑛𝑠
𝑠𝑑 = 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑜𝑓 𝑡ℎ𝑒 𝑑𝑖𝑓𝑓𝑒𝑟𝑒𝑛𝑐𝑒𝑠
6. State the decision rule.

- If the computed value of the statistic is greater than the critical value, then reject the null
hypothesis.
-If the computed value of the statistic is less than the tabular or critical value, then do not
reject the null hypothesis.

7. Conclusion
- Based on the test results, the researcher will reach a conclusion about the population
under study.

Let’s get started!

Example 1.
To determine whether the students’ performance in College Algebra will improve after enrolling
in the subject for one term, a 60-item pre-test and post-test are administered to them on the first
day and last day of classes respectively. Use 1% level of significance.

The results are as follows:

Student Pre- test Post- test Difference (d) d2


A 34 45 -11 121
B 23 32 -9 81
C 40 46 -6 36

PROBABILITY AND STATISTICS P a g e 5 | 11


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
D 31 57 -26 676
E 24 39 -15 225
F 45 48 -3 9
G 27 27 0 0
H 32 33 1 1
I 12 18 6 36
J 45 45 0 0
Σ𝑑 = −77 2
Σ𝑑 = 1185

Solution:

Ho: Students’ performance in College Algebra did not improve.


a. State the null and 𝜇𝑑 = 0
alternative hypotheses. Ha: Students’ performance in College Algebra did improve.
𝜇𝑑 < 0

Since it has been stated from the problem that the significant level
b. Select the level of is 1% therefore,
significance
α = 0.01

c. Identify the appropriate one-tailed test


test-statistic and paired sample t-test
determine whether it is
one-tailed or two-tailed.

Use the t-distribution table since our test-statistic is a paired


sample t-test

α = 0.05, 𝑑𝑓 = 𝑛 − 1 = 10 − 1 = 9
d. Determine the critical
value and the rejection Refer on the table α = 0.01,
region/s. 𝑑𝑓 = 9, one-tailed test

The critical value is -2.821

e. Compute the test Given: 𝑑 = −77 𝑛 = 10


statistic.
𝑛(Σ𝑑 2 ) − (Σ𝑑 )2 10(1185) − (−77)2
𝑠𝑑 = √ = √
𝑛(𝑛 − 1) 10(10 − 1)

11850 − 5929 5921


𝑠𝑑 = √ =√ = √65.79 = 8.11
10(9) 90

PROBABILITY AND STATISTICS P a g e 6 | 11


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Substitute.
Σ𝑑 −77
− 𝜇 𝑑 −0 −.7.7
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = 𝑛𝑠 = 10 = = 𝟑. 𝟏𝟔
𝑑 8.11 8.11
√𝑛 √10 3.33

f. State the decision rule. Since the t-computed (3.16) is greater than the t-tabular value (-
2.821), we reject the null hypothesis (Ho).

g. Conclusion The performance of the students in College Algebra has


significantly improved.

Example 2.
A study was conducted to determine the effectiveness of a weight loss program. The table below
shows the before and after weight of 12 subjects in the program. Is this program effective for
reducing weight? (Use a 5% significance level

S BEFORE AFTER Difference SD


1 185 169 16 256
2 192 187 5 25
3 206 193 13 169
4 177 176 1 1
5 225 194 31 961
6 168 171 -3 9
7 256 228 28 784
8 239 217 22 484
9 199 204 -5 25
10 218 195 23 529
11 200 201 -1 1
12 190 188 2 4
Σ𝑑 = 132 2
Σ𝑑 = 3248

Solution:

a. State the null and Ho: The weight loss program is effective.
alternative hypotheses. 𝜇𝑑 ≥ 0
Ha: The weight loss program is not effective.
𝜇𝑑 < 0
b. Select the level of The significance level was not mentioned in the problem. In this
significance case, we use the 5%.

PROBABILITY AND STATISTICS P a g e 7 | 11


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
α = 0.05

c. Identify the appropriate since n=12 and 12 is less than 30 and the value of the population
test-statistic and mean is not equal to the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = paired sample t-test (one - tailed)

d. Determine the critical We use the t-distribution table since our test-statistic is paired
value and the rejection sample t-test.
region/s.
Given, α = 0.05, 𝑑𝑓 = 𝑛 − 1 = 12 − 1 = 11

Refer on the table α = 0.05,


𝑑𝑓 = 11, one-tailed test

The critical value is -1.796

e. Compute the test Solving for the standard deviation;


statistic.
𝑛(Σ𝑑 2 ) − (Σ𝑑 )2 12(3248) − (132)2
𝑠𝑑 = √ =√
𝑛(𝑛 − 1) 12(12 − 1)

38976 − 17424 21552


𝑠𝑑 = √ =√ = √163.27 = 𝟏𝟐. 𝟕𝟖
12(11) 132

Therefore,
𝒔𝒅 = 𝟏𝟐. 𝟕𝟖

Σ𝑑 = 132 𝑛 = 12

Substitute.
Σ𝑑 132
− 𝜇𝑑
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = 𝑛 = 12 − 0 = 11 = 𝟐. 𝟗𝟖
𝑠𝑑 12.78 3.69
√𝑛 √12

f. State the decision rule. Since t-computed (2.98) is greater than the t-tabular value (1.796),
we reject the null hypothesis (Ho).

g. Conclusion There is enough evidence to prove that the weight loss program is
effective.

PROBABILITY AND STATISTICS P a g e 8 | 11


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Example 3.
A dietitian wishes to see if a person’s cholesterol level will change if the diet is supplemented by a
certain mineral. Six subjects were pretested, and then they took the mineral supplement for a 6-
week period. The results are shown in the table. Can it be concluded that the cholesterol level has
been changed at 𝛼 = 0.01?

Subject 1 2 3 4 5 6
Before (x1) 210 235 208 190 172 244
After (x2) 190 170 210 188 173 228

Solution:
a. State the null and Ho: The person’s cholesterol level will not change if the diet is
alternative hypotheses. supplemented by a certain mineral.
𝜇𝑑 = 0
Ha: The person’s cholesterol level will change if the diet is
supplemented by a certain mineral.
𝜇𝑑 ≠ 0
b. Select the level of Since it has been stated from the problem that the significant level
significance is 0.01 therefore,

α = 0.01

c. Identify the appropriate since n=6 and 6 is less than 30 and the value of the population
test-statistic and mean is not equal to the specified value, therefore,
determine whether it is
one-tailed or two-tailed. test statistic = paired sample t-test (two-tailed)

d. Determine the critical We use the t-distribution table since our test-statistic is paired
value and the rejection sample t-test.
region/s.
Given, α = 0.01

𝑑𝑓 = 𝑛 − 1 = 6 − 1 = 5
Refer on the table α = 0.01,
𝑑𝑓 = 5, two-tailed test

Critical value = ±4.032 (±, since it’s two tailed, the rejection region
is found at both ends of the normal curve)

e. Compute the test


statistic.

PROBABILITY AND STATISTICS P a g e 9 | 11


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Subject Before After d d2
1 210 190 20 400
2 235 170 65 4225
3 208 210 -2 4
4 190 188 2 4
5 172 173 -1 1
6 244 228 16 256
Σ𝑑 =
Σ𝑑 2 = 4890
100
Solving for the standard deviation;

𝑛(Σ𝑑 2 ) − (Σ𝑑 )2 6(4890) − (100)2


𝑠𝑑 = √ =√
𝑛(𝑛 − 1) 6(6 − 1)

29340 − 10000 19340


=√ =√ = √644.67 = 𝟐𝟓. 𝟑𝟗
6(5) 30

𝑑 = 100 𝑛=6

Substitute.
Σ𝑑 100
− 𝜇 𝑑 16.67
𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 = 𝑛𝑠 = 6 = = 𝟏. 𝟔𝟎𝟕
𝑑 25.39 10.37
√𝑛 √6
f. State the decision rule. Since t-computed 1.607 is less than the t-tabular value ±4.032, we
do not reject the null hypothesis (Ho) .

g. Conclusion Therefore, there is not enough evidence that the person’s


cholesterol level will change if the diet is supplemented by a certain
mineral.

Generalizations.
Before we finally end this chapter, we summarize the ideas and the pieces of information with this
week’s lesson.
The paired sample t- test is a statistical hypothesis used to test a difference between means for
dependent samples. In testing hypothesis, follow the steps in hypothesis testing. First, state the
hypothesis and identify the claim. Remember that paired sample t-test is a statistical procedure
used to determine whether the mean difference between two sets of observations is zero.
Therefore, your Ho would always be 𝜇𝑑 = 0 , which means that there is no difference or there is
no change between the two sets of test. Then, select the level of significance, identify the
appropriate test-statistic (paired sample t-test) and determine whether it is one-tailed or two-

PROBABILITY AND STATISTICS P a g e 10 | 11


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
tailed, determine the critical value and the rejection region/s, compute the test statistic and then
decide and conclude.

For this week’s lesson, it helps us realize that in life, there is always a chance of improvement and
a place for change.

PROBABILITY AND STATISTICS P a g e 11 | 11


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Final Term – Second Semester Week 6: May 3-7, 2021

I. INTRODUCTION

Job well done Louisian GEMs. We are near to the end of the line. Keep the spirit and continue
soaring high for the remaining weeks of this school year.

This week, you shall be given another lesson to study and enrichment activities to solve. Attached
to this module is the weekly Study and Assessment Guide

DATE TOPIC ACTIVITIES OR TASKS


One-Way Analysis of Variance Read the lessons.
May 3 - 7, 2021
Assessment 5

For these weeks of the final term, the following shall be your guide for the different lessons and
tasks that you need to accomplish. Be patient, read it carefully before proceeding to the tasks expected of
you. GOOD LUCK!

Content One-Way Analysis of Variance


Learning Competencies The learner should be able to:
• define the analysis of variance;
• compute the F – value;
• perform the steps of hypothesis testing of the Analysis of
Variance.
Activities Assessment 5
Essential Questions How do we define analysis of variance?
When do we apply F-test?
Value Statement How would we compare more than 2 means?
References Textbooks:

Alferez, M. S. & Duro, M. C. A. (2018). Statistics and Probability. MSA


Publishing House. Cainta, Philippines.

De Guzman, D. (2017). Statistics and Probability. C & E Publishing Inc.


Quezon City, Philippines

Albert, J. R., Albacea, Z., Ayaay, M. J., David, I., & de Mesa, I. (2016)
Statistics and Probability. Commission on Higher Education. Manila,
Philippines

Statistics and Probability P a g e 1 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
II. LEARNING CONTENT

A researcher might, for example, test students from multiple colleges to see if students from one of the
colleges consistently outperform students from the other colleges.

What statistical test will he use in his study? Another thing, for example, we might ask whether the
difference between two sample means could have been produced by chance.

What if our experiment had more than two conditions or groups?


We would have more than 2 means. We would have one mean for each group or condition. That
could be a lot depending on the experiment.

How would we compare all of those means? What should we do, run a lot of t-tests, comparing every
possible combination of means?

Actually, you could do that. Or, you could do an ANOVA.

Analysis of Variance commonly known as ANOVA is a technique that uses the F test to test a hypothesis
concerning the means of three (3) or more populations.

The t- and z-test methods developed in the 20th century were used for statistical analysis until 1918, when
Sir Ronald Fisher created the analysis of variance method. ANOVA is also called the Fisher analysis of
variance, and it is the extension of the t- and z-tests. The term became well-known in 1925, after
appearing in Fisher's book, "Statistical Methods for Research Workers. It was employed in experimental
psychology and later expanded to subjects that were more complex.

Two different estimates of the population variance are made. The first estimate is called the between-
group variance. It involves computing the variance by using the means of the groups or between the
groups. The second estimate, the within-group variance, is made by computing the variance using all the
data. It is not affected by differences in the means.

For a test of difference among three or more means, the following hypotheses are used:
Ho: μ1 = μ2 = μ3 = … = μn
Ha: At least one mean is different from the others

If there is no difference in the means, the between – group variance estimate will approximately be equal
to the within-group variance estimate, and the F test value will be approximately equal to 1. When the
means differ significantly, the between-group variance will be much larger than the within-group variance.
Then the F test value will be significantly greater than 1. Then the null hypothesis will be rejected.

The degree of freedom (df) for the F test are d.f.N = k - 1 where k is the number of groups and
d.f.D = n - k where n is the sum of the sample sizes of the groups, n = n 1+n2+…+nk. The sample sizes need
not be equal. The F test to compare the means is always a right-tailed test.

Here are the steps in computing the F-test:


A. Compute the ∑ 𝑥 2 , square all values include in the different samples and then add.

B. Solve for the between – samples sum of squares (SSB).

Statistics and Probability P a g e 2 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
We need first to get the total values per sample and the total number of values per sample. Then,
get the grand total of all the values and the total number of all the samples.

The between-samples sum of squares, denoted by SSB,


2
𝑇12 𝑇22 𝑇32 (∑ 𝑥)
𝑆𝑆𝐵 = ( + + ⋯+ ) −
𝑛1 𝑛2 𝑛𝑘 𝑛
where Ti = 1, 2, …, k = sum of sample in the ith group
ni = 1, 2, …, k = number of samples in the ith group

C. Solve for the within-samples sum of squares (SSW).


The within-samples sum of squares, denoted by SSW,

𝑇12 𝑇22 𝑇𝑘2


𝑆𝑆𝑊 = ∑ 𝑥 2 − ( + + ⋯+ )
𝑛1 𝑛2 𝑛𝑘

D. Solve for the variance between samples (MSB) and the variance within samples (MSW).
𝑆𝑆𝐵 𝑆𝑆𝑊
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
𝑘−1 𝑛−𝑘
where k – 1 and n – k are respectively, the degrees of freedoms for the numerator and
denominator for the F distribution.

E. Compute the F test statistic.


The value of the test statistic F for an ANOVA test is calculated as
𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒 𝑏𝑒𝑡𝑤𝑒𝑒𝑛 𝑠𝑎𝑚𝑝𝑙𝑒𝑠 𝑀𝑆𝐵
𝐹= =
𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒 𝑤𝑖𝑡ℎ𝑖𝑛 𝑠𝑎𝑚𝑝𝑙𝑒𝑠 𝑀𝑆𝑊

F. After solving all the needed computations, input all computations on the ANOVA Table

Calculations of F test are often recorded in a table called ANOVA table.


Sources of Degrees of Sum of Mean F-test
Variation Freedom Squares Square Statistic
Between k–1 SSB MSB 𝑀𝑆𝐵
Within n-k SSW MSW 𝑀𝑆𝑊
Total n-1 SST

The sum of SSB and SSW is called the total sum of squares and it is denoted by SST. That is, SST = SSB
+SSW

When the null hypothesis is rejected in ANOVA, it only means that at least one mean is different from the
others. To locate the difference or differences among the means, it is necessary to use other tests such as
the Tukey or the Scheffé test.

Use the following table for the critical value depending on the level of significance, degree of freedom of
Numerator (d.f.N.) and degree of freedom of Denominator (d.f.D.)

Statistics and Probability P a g e 3 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Statistics and Probability P a g e 4 | 17
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Statistics and Probability P a g e 5 | 17
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Let us solve some examples:

1. An introductory Calculus course is taken by students with varying high school records. The values
given are the final numerical averages in the Calculus course. At α = 0.05, test the claim that the
mean scores are equal in the three groups.

A B C
90 80 60
86 70 60
88 61 55
93 52 62
80 73 50
65 70
83

Solution:
Use the steps of the hypothesis testing together with the steps of solving the F test.

1. State the null and alternative hypothesis.


• Ho: μ1 = μ2 = μ3 (claim)
• Ha: At least one mean is different from the others

2. Select the level of significance.


• Based on the problem, the level of significance is at 0.05.

3. Identify the test-statistic.


• Based on the data given, there are 3 groups which is also 3 means involve so, F -test
is the statistical test of the problem.

4. Determine the critical value and the rejection region/s.


d.f.N = k – 1, where k is the number of groups
d.f.D = n – k, where n is the sum of the sample sizes of the groups, n=n 1+n2+…+nk
Note: The sample sizes need not be equal.
k=3 n = 18

d.f.N = k – 1 = 3 – 1 = 2 d.f.D = n – k = 18 – 3 = 15

Use the F distribution table. Consider the intersection of d.f.N (degree of freedom of
Numerator), d.f.D (degree of freedom of Denominator) at the level of significance which
0.05.

Statistics and Probability P a g e 6 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Critical value is 3.68

5. Compute the test statistic.


Here is how you will solve the F-test.
A. Compute the ∑ 𝑥 2 , square all values include in the different samples and then add.
∑ 𝑥 2 = 902 + 862 + 882 + 932 + 802 + 802 + 702 + 612 + 522 + 732 + 652 + 832 + 602
+ 602 + 552 + 622 + 502 + 702
∑ 𝑥 2 = 93,926

B. Solve for the between – samples sum of squares (SSB). We need first to get the total values per
sample and the total number of values per sample.
Total of A: 437; n is 5
Total of B: 484; n is 7
Total of C: 357, n is 6

2
𝑇12 𝑇22 𝑇𝑘2 (∑ 𝑥)
𝑆𝑆𝐵 = ( + +⋯+ )−
𝑛1 𝑛2 𝑛𝑘 𝑛

4372 4842 3572 12782


𝑆𝑆𝐵 = ( + + )−
5 7 6 18
𝑆𝑆𝐵 = 2162.44

Statistics and Probability P a g e 7 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
C. Solve for the within-samples sum of squares (SSW).
𝑇12 𝑇22 𝑇𝑘2
𝑆𝑆𝑊 = ∑ 𝑥 2 − ( + + ⋯+ )
𝑛1 𝑛2 𝑛𝑘

4372 4842 3572


𝑆𝑆𝑊 = 93926 − ( + + )
5 7 6
𝑆𝑆𝑊 = 1025.56

D. Solve for the variance between samples (MSB) and the variance within samples (MSW).
𝑆𝑆𝐵 𝑆𝑆𝑊
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
𝑘−1 𝑛−𝑘
2162.44 1025.56
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
3−1 18 − 3
𝑀𝑆𝐵 = 1081.22 𝑀𝑆𝑊 = 68.37

E. Compute the F-value.


𝑀𝑆𝐵 1081.22
𝐹= = = 𝟏𝟓. 𝟖𝟏
𝑀𝑆𝑊 68.37
F. Here is the ANOVA table:

Sources of Degrees of Sum of Mean F-test


Variation Freedom Squares Square Statistic
Between 2 2162.44 1081.22
Within 15 1025.56 68.37 15.81
Total 17 3188

6. State the decision rule.


If the 𝐹𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 > 3.68, Reject Ho.
Since the 𝐹𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑖𝑠 15.81, 𝑡ℎ𝑒𝑛 15.81 > 3.68, the decision is Reject Ho.

7. State the conclusion.


There is enough evidence to reject the claim and conclude that at least one mean is different
from the others.

2. A study was done before the recent surge in gasoline prices then to compare the cost to drive
25 miles for different types of hybrid vehicles. The cost of a gallon of gas at the time of the
study was approximately $2.50. Based on the information given below for different models
of hybrid cars, trucks, and SUVs, is there sufficient evidence to conclude a difference in the
mean cost to drive 25 miles? Use α = 0.05.
Hybrid Cars Hybrid SUVs Hybrid Trucks
2.10 2.10 3.62
2.70 2.42 3.43
1.67 2.25
1.67 2.10
1.30 2.25

Statistics and Probability P a g e 8 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Solution:
Use the steps of the hypothesis testing together with the steps of solving the F test.

1. State the null and alternative hypothesis.


• Ho: μ1 = μ2 = μ3
• Ha: At least one mean is different from the others (Claim: There is a sufficient
evidence to conclude a difference in the mean cost to drive 25 miles.)

2. Select the level of significance.


• Based on the problem, the level of significance is at 0.05.

3. Identify the test-statistic.


• Based on the data given, there are 3 groups which is also 3 means involve so, F -test
is the statistical test of the problem.

4. Determine the critical value and the rejection region/s.


d.f.N = k – 1, where k is the number of groups
d.f.D = n – k, where n is the sum of the sample sizes of the groups, n=n1+n2+…+nk
Note: The sample sizes need not be equal.
k=3 n = 12

d.f.N = k – 1 = 3 – 1 = 2 d.f.D = n – k = 12 – 3 = 9

Use the F distribution table. Consider the intersection of d.f.N (degree of freedom of
Numerator), d.f.D (degree of freedom of Denominator) at the level of significance which
0.05.

Statistics and Probability P a g e 9 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Critical value is 4.26

5. Compute the test statistic.


Here is how you will solve the F-test.

A. Compute the ∑ 𝑥 2 , square all values include in the different samples and then add.
∑ 𝑥 2 = 2.102 + 2.702 + 1.672 + 1.672 + 1.302 + 2.102 + 2.422 + 2.252 + 2.102 + 2.252
+ 3.622 + 3.432
∑ 𝑥 2 = 68.64

B. Solve for the between – samples sum of squares (SSB). We need first to get the total values per
sample and the total number of values per sample.
Total of Cars: 9.44; n is 5
Total of SUVs: 11.12; n is 5
Total of Trucks: 7.05, n is 2

2
𝑇12 𝑇22 𝑇𝑘2 (∑ 𝑥)
𝑆𝑆𝐵 = ( + +⋯+ )−
𝑛1 𝑛2 𝑛𝑘 𝑛

9.442 11.122 7.052 27.612


𝑆𝑆𝐵 = ( + + )−
5 5 2 12
𝑆𝑆𝐵 = 3.88

Statistics and Probability P a g e 10 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
C. Solve for the within-samples sum of squares (SSW).

𝑇12 𝑇22 𝑇𝑘2


𝑆𝑆𝑊 = ∑ 𝑥 2 − ( + + ⋯+ )
𝑛1 𝑛2 𝑛𝑘
9.442 11.122 7.052
𝑆𝑆𝑊 = 68.64 − ( + + )
5 5 2
𝑆𝑆𝑊 = 1.24

D. Solve for the variance between samples (MSB) and the variance within samples (MSW).
𝑆𝑆𝐵 𝑆𝑆𝑊
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
𝑘−1 𝑛−𝑘
3.88 1.24
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
3−1 12 − 3
𝑀𝑆𝐵 = 1.94 𝑀𝑆𝑊 = 0.14

E. Compute the F-value.


𝑀𝑆𝐵 1.94
𝐹= = = 𝟏𝟑. 𝟖𝟔
𝑀𝑆𝑊 0.14
F. Here is the ANOVA table:

Sources of Degrees of Sum of Mean F-test


Variation Freedom Squares Square Statistic
Between 2 3.88 1.94
Within 9 1.24 0.14 13.86
Total 11 5.12

6. State the decision rule.


Reject Ho, if the 𝑓𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 > 4.26
Since the 𝑓𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑖𝑠 13.86, 𝑡ℎ𝑒𝑛 13.86 > 3.68, the decision is Reject Ho.

7. State the conclusion.


There is a sufficient evidence to conclude a difference in the mean cost to drive 25 miles.

3. The manager of a bank decided to compare the speed on the job of four tellers working in the
bank. The following data give the amount of time (in minutes) that they spent serving their
customers, picked at random.

Teller 1 Teller 2 Teller 3 Teller 4


0.8 8.7 9.7 0.7
1.6 4.2 2.0 3.2
8.8 0.2 5.9 4.2
7.7 9.0 2.3 6.7
3.3 5.8 7.5
1.4

Statistics and Probability P a g e 11 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Is there evidence indicating that there are significant differences among the true mean serving time for
the four tellers? Use α = 0.025

Solution:
Use the steps of the hypothesis testing together with the steps of solving the F test.

1. State the null and alternative hypothesis.


• Ho: μ1 = μ2 = μ3
• Ha: At least one mean is different from the others (Claim: There are significant
differences among the true mean serving time for the four tellers.)

2. Select the level of significance.


• Based on the problem, the level of significance is at 0.025.

3. Identify the test-statistic.


• Based on the data given, there are 3 groups which is also 3 means involve so, F -test
is the statistical test of the problem.

4. Determine the critical value and the rejection region/s.


d.f.N = k – 1, where k is the number of groups
d.f.D = n – k, where n is the sum of the sample sizes of the groups, n=n 1+n2+…+nk
Note: The sample sizes need not be equal.
k=4 n = 20

d.f.N = k – 1 = 4 – 1 = 3 d.f.D = n – k = 20 – 4 = 16

Use the F distribution table. Consider the intersection of d.f.N (degree of freedom of
Numerator), d.f.D (degree of freedom of Denominator) at the level of significance which
0.025.

Statistics and Probability P a g e 12 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Critical value is 4.08

5. Compute the test statistic.


Here is how you will solve the F-test.
A. Compute the ∑ 𝑥 2 , square all values include in the different samples and then add.
∑ 𝑥 2 = 0.82 + 1.62 + 8.82 + 7.72 + 3.32 + 8.72 + 4.22 + 0.22 + 9.02 + 9.72 + 2.02 + 5.92
+ 2.32 + 5.82 + 0.72 + 3.22 + 4.22 + 6.72 + 7.52 + 1.42
∑ 𝑥 2 = 628.49

B. Solve for the between – samples sum of squares (SSB). We need first to get the total values per
sample and the total number of values per sample.
Teller 1: 22.2; n is 5
Teller 2: 22.1; n is 4
Teller 3: 25.7, n is 5
Teller 4: 23.7, n is 6

2
𝑇12 𝑇22 𝑇𝑘2 (∑ 𝑥)
𝑆𝑆𝐵 = ( + +⋯+ )−
𝑛1 𝑛2 𝑛𝑘 𝑛

22.22 22.12 25.72 23.72 93.72


𝑆𝑆𝐵 = ( + + + )−
5 4 5 6 20
𝑆𝑆𝐵 = 7.40

Statistics and Probability P a g e 13 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
C. Solve for the within-samples sum of squares (SSW).

𝑇12 𝑇22 𝑇𝑘2


𝑆𝑆𝑊 = ∑ 𝑥 2 − ( + + ⋯+ )
𝑛1 𝑛2 𝑛𝑘
22.22 22.12 25.72 23.72
𝑆𝑆𝑊 = 628.49 − ( + + + )
5 4 5 6
𝑆𝑆𝑊 = 182.11

D. Solve for the variance between samples (MSB) and the variance within samples (MSW).
𝑆𝑆𝐵 𝑆𝑆𝑊
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
𝑘−1 𝑛−𝑘
7.40 182.11
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
4−1 20 − 4
𝑀𝑆𝐵 = 2.47 𝑀𝑆𝑊 = 11.38

E. Compute the F-value.


𝑀𝑆𝐵 2.47
𝐹= = = 𝟎. 𝟐𝟐
𝑀𝑆𝑊 11.38
F. Here is the ANOVA table:

Sources of Degrees of Sum of Mean F-test


Variation Freedom Squares Square Statistic
Between 3 7.40 2.47
Within 16 182.11 11.38 0.22
Total 19 189.51

6. State the decision rule.


If the 𝐹𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 < 4.08. do not reject Ho.
Since the 𝐹𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑖𝑠 0.22, 𝑡ℎ𝑒𝑛 0.22 < 4.08, the decision is do not reject Ho.

7. State the conclusion.


There is no significant difference among the true mean serving time for the four tellers.

4. A pharmacology lab conducted an experiment to compare four pain-relieving drugs. 24 subjects


were used. The following figures give the number of hours of pain relief provided subsequent to
administering a drug.
Drug A Drug B Drug C Drug D
8 9 6 5
7 1 9 7
4 8 5 2
3 4 3 8
5 9 2 9
3 8 2
1
3

Statistics and Probability P a g e 14 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
At the 5% level of significance, are the drugs different in the true mean number of hours of relief
provided?

Solution:
Use the steps of the hypothesis testing together with the steps of solving the F test.

1. State the null and alternative hypothesis.


• Ho: μ1 = μ2 = μ3
• Ha: At least one mean is different from the others (Claim: The drugs differ in the true
mean number of hours of relief provided.)

2. Select the level of significance.


• Based on the problem, the level of significance is at 0.05.

3. Identify the test-statistic.


• Based on the data given, there are 3 groups which is also 3 means involve so, F -test
is the statistical test of the problem.

4. Determine the critical value and the rejection region/s.


d.f.N = k – 1, where k is the number of groups
d.f.D = n – k, where n is the sum of the sample sizes of the groups, n=n 1+n2+…+nk
Note: The sample sizes need not be equal.
k=4 n = 25

d.f.N = k – 1 = 4 – 1 = 3 d.f.D = n – k = 25 – 4 = 21

Use the F distribution table. Consider the intersection of d.f.N (degree of freedom of
Numerator), d.f.D (degree of freedom of Denominator) at the level of significance which
0.025.

Statistics and Probability P a g e 15 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Critical value is 3.07

5. Compute the test statistic.


Here is how you will solve the F-test.
G. Compute the ∑ 𝑥 2 , square all values include in the different samples and then add.
∑ 𝑥 2 = 82 + 72 + 42 + 32 + 52 + 32 + 92 + 12 + 82 + 42 + 92 + 82 + 12 + 32 + 62 + 92
+ 52 + 32 + 22 + 52 + 72 + 22 + 82 + 92 + 22
∑ 𝑥 2 = 871

H. Solve for the between – samples sum of squares (SSB). We need first to get the total values per
sample and the total number of values per sample.
Drug A: 30; n is 6
Drug B: 43; n is 8
Drug C: 25, n is 5
Drug D: 33, n is 6

2
𝑇12 𝑇22 𝑇𝑘2 (∑ 𝑥)
𝑆𝑆𝐵 = ( + +⋯+ )−
𝑛1 𝑛2 𝑛𝑘 𝑛

302 432 252 332 1312


𝑆𝑆𝐵 = ( + + + )−
6 8 5 6 25
𝑆𝑆𝐵 = 1.19

Statistics and Probability P a g e 16 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
I. Solve for the within-samples sum of squares (SSW).

𝑇12 𝑇22 𝑇𝑘2


𝑆𝑆𝑊 = ∑ 𝑥 2 − ( + + ⋯+ )
𝑛1 𝑛2 𝑛𝑘

302 432 252 332


𝑆𝑆𝑊 = 871 − ( + + + )
6 8 5 6
𝑆𝑆𝑊 = 183.38

J. Solve for the variance between samples (MSB) and the variance within samples (MSW).
𝑆𝑆𝐵 𝑆𝑆𝑊
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
𝑘−1 𝑛−𝑘
1.19 183.38
𝑀𝑆𝐵 = 𝑀𝑆𝑊 =
4−1 25 − 4
𝑀𝑆𝐵 = 0.40 𝑀𝑆𝑊 = 8.73

K. Compute the f-value.


𝑀𝑆𝐵 0.40
𝐹= = = 𝟎. 𝟎𝟓
𝑀𝑆𝑊 8.73
L. Here is the ANOVA table:

Sources of Degrees of Sum of Mean F-test


Variation Freedom Squares Square Statistic
Between 3 1.19 0.40
Within 21 183.38 8.73 0.05
Total 24 184.57

6. State the decision rule.


If the 𝑓𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 < 3.07. do not reject Ho.
Since the 𝑓𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 𝑖𝑠 0.05, 𝑡ℎ𝑒𝑛 0.05 < 4.08, the decision is do not reject Ho.

7. State the conclusion.


The drugs do not differ in the true mean number of hours of relief provided.

Generalization:

We can use F-test to test the difference of more than 2 means. Follow the procedures in getting the F-
value and the hypothesis testing so that you can give a sound conclusion to a problem or study.

Note: If you have questions regarding your lessons, feel free to message your subject teacher.

Statistics and Probability P a g e 17 | 17


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Finals – Second Semester Week 7: May 10-14, 2021

I. INTRODUCTION
This week, you shall be given another lesson to study and enrichment activities to solve.
Attached to this 7th week module is the weekly Study and Assessment Guide

DATE TOPIC ACTIVITIES OR TASKS


Correlation Analysis: Read and analyze the lessons.
Pearson Product-Moment Correlation
May 10-14,2021
Coefficient Assessment 6
Regression Analysis

For this week of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the
tasks expected of you. GOOD LUCK!

Content Correlation Analysis:


Pearson Product-Moment Correlation Coefficient
Regression Analysis
Learning Competencies • Define correlation analysis, scatter diagram and Pearson-r
correlation
• Know how and when to use the Pearson-r correlation
• Identify the different scatter diagram
• Solve for the computed value of Pearson-r correlation
• Solve and test the significance of r on the given problem
• Compute the equation of linear regression
Activities • Assessment 6
Essential Question/s • Is relationship mutual?
• How do we make predictions and decisions based on
current numerical information?
Value Statement • “ In the name of freedom, there has to be a correlation
between rights and duties, by which every person is called
to assume responsibility for his or her choices, made as a
consequence of entering into relations with others.” -
Pope Benedict XVI
References Textbooks:
Rumsey. D (2021), Statistics Workbook For Dummies, Statistics II
For Dummies, and Probability For Dummies.

Myers, R., Walpole, R., [Link]. (2017). Probability & Statistics for
Engineers & Scientists, Pearson Education Limited, England.

Alferez, M. S. & Duro, M. C. A. (2018). Statistics and Probability.

Statistics and Probability P a g e 1 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
MSA Publishing House. Cainta, Philippines.

De Guzman, D. (2017). Statistics and Probability. C & E Publishing


Inc. Quezon City, Philippines

Alferez, M. S. & Duro, M. C. A. (2018). Statistics and Probability.


MSA Publishing House. Cainta, Philippines.

I. LEARNING CONTENT

The main objective of many statistical investigation is to establish relationship between two
variables. For example: A health researcher wants to determine the relationship between a
mother’s weight and her baby’s birth weight, An educational researcher wants to correlate the
grade of his students in science and math, A psychologist may want to know the relationship
between the IQ and personality of Individual.

These are just some examples of the situations that we can apply correlational test.

Correlation Analysis is the process of exploring the relationship or association between


variables. The degree of association is measured by a correlation coefficient, denoted by r. This
also defined as the measure of the linear relationship between two random variable x and y.

➢ This test also attempts to measure the strength of such relationships between two
variables namely (x and y) by the means of single number called a correlation coefficient.
➢ The variables x and y are referred to as bivariate variable.

Among the wide classes of correlational test, The Pearson Product Moment Coefficient of
Correlation (Pearson-r) and the Spearman Rank-Order Coefficient of Correlation are the two
most commonly used.

Correlation Coefficient
A correlation coefficient is a number between -1 and +1 which provides a measure of
the strength or degree of the linear association between two variables, x and y. The + sign and –
sign indicate only the direction of the correlation.
Remember: + is for a positive correlation and - is for a negative correlation.

There are 3 degrees of relationship or correlation between two variables


1. Perfect correlation (Positive and Negative)
2. Some degree of correlation (Positive and Negative)
3. No correlation.

Statistics and Probability P a g e 2 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
A positive correlation implies a direct relationship between the variables. The two variables
increase or decrease proportionately. If one variable increases in the same proportion while the
other variable decreases and vice versa, the two variables are said to be negatively correlated.
The quantitative interpretation of the degree of the linear relationship existing is shown in the
following range.

±1.00 Perfect positive or negative correlation

±0.91 - ±0.99 Very high positive or negative correlation

±0.71 - ±0.90 High positive or negative correlation

±0.51 - ±0.70 Moderately positive or negative correlation

±0.31 - ±0.50 Low positive or negative correlation

±0.01 - ±0.30 negligible positive or negative correlation

0.00 No correlation

For example, if the result is 0.90 it indicates that it has a very high positive correlation but if the
result is -0.90 it indicates that it has a very high negative correlation. The signs “ + and – “
indicate the direction.
Scatter Diagram (Scatterpoint Diagram)
➢ This diagram is used to illustrate the relationship between two variables
➢ It is a graphic visualization of the relationship between the bivariate data.

Statistics and Probability P a g e 3 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Figure a illustrates the two variables which are both increasing. The data are forming a straight
line with a positive slope making it a perfect positive correlation.
However, Figure d illustrates a negative slope which also means that one variable increase while
the other decreases.
Figure b shows a positive correlation because the data points are forming a line from left to
right.
Figure c shows no relationship at all because no trend is being seen as all data point are literally
scattered in the diagram.

Statistics and Probability P a g e 4 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
The Pearson Product- Moment Correlation Coefficient (Pearson r)
Sir Karl Pearson developed a rigorous mathematical treatment to describe the relationship
between two variables now known as the Person Product- Moment Coefficient of Correlation,
denoted by r. The size of the correlation varies from +1 through 0 to -1.

The formula for Pearson r is

𝑛 ∑ 𝑋𝑌 − (∑ 𝑋)( ∑ 𝑌) 𝑋 is the first variable under study


𝑟=
√[𝑛 ∑ 𝑋 2 − (∑ 𝑋)2 ][𝑛 ∑ 𝑌 2 − (∑ 𝑌)2 ] 𝑌 is the first variable under study
𝑛 is the sample size

Steps in Solving for Pearson-r:


1. Find the summation of the variables (X and Y)
2. Multiply each pair of variable ( X and Y), then get the summation of the product.
3. Square each variable (X and Y) then get the summation of each variable.
4. Compute for the value of Pearson-r.
5. State the Conclusion

Testing the Significance of 𝑟


The correlation coefficient r is an estimate of the population correlation coefficient in the same
sense that 𝑥̅ is an estimate of the population mean 𝜇. This is actually the part where we will
make our decision.
The formula is
𝑟 is the degree of relationship between the two variables
𝑛−2
𝑡 = 𝑟 (√ ) 𝑛 is the sample size
1 − 𝑟2 Degrees of freedom: df = n - 2

Steps in Testing the Significance of r.


1. State the null and alternative hypothesis
2. Determine the level of significance and the degree of freedom
3. Using the level of significance and degree of freedom, locate for the critical values of t
on the T-Distribution table, two-tailed (this was discussed on week 1)
Note: When testing the significant relationship of two variables, we use a two tailed test.
4. Compute for the value of the t-statistics.
5. Make a decision.
6. State the Conclusion.

Statistics and Probability P a g e 5 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Example 1. A statistic teacher recorded the Math and Science test scores of twelve senior high
students from one section. Using 0.05 as the level of significance, is there a significant
relationship between the test scores in Mathematics and Science?

Student Math Scores Science Scores


Number (X) (Y)
1 18 20
2 16 18
3 11 12
4 15 17
5 15 15
6 11 14
7 11 12
8 13 14
9 8 10
10 9 13
11 13 12
12 7 9

Let’s first get the value of 𝑟.


1st Step:
Find the summation of X and Y. Get the sum of all the Math scores ∑ 𝑋 and Science scores ∑ 𝑌.
Student Number Math Scores Science Scores
(X) (Y)
1 18 20
2 16 18
3 11 12
4 15 17
5 15 15
6 11 14
7 11 12
8 13 14
9 8 10
10 9 13
11 13 12
12 7 9
N=12 ∑ 𝑋 = 147 ∑ 𝑌 = 166

Statistics and Probability P a g e 6 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
2nd Step: Get the product of each pair of X and Y entries. Get the total of XY.

Student Math Scores Science Scores


XY
Number (X) (Y)
1 18 20 360
2 16 18 288
3 11 12 132
4 15 17 255
5 15 15 225
6 11 14 154
7 11 12 132
8 13 14 182
9 8 10 80
10 9 13 117
11 13 12 156
12 7 9 63
N=12 ∑ 𝑋 = 147 ∑ 𝑌 = 166 ∑ 𝑋𝑌 = 2144

3rd Step: Square each variable ( X and Y ) then compute for its summation

Science
Student Math Scores
Scores XY X2 Y2
Number (X)
(Y)
1 18 20 360 324 400
2 16 18 288 256 324
3 11 12 132 121 144
4 15 17 255 225 289
5 15 15 225 225 225
6 11 14 154 121 196
7 11 12 132 121 144
8 13 14 182 169 196
9 8 10 80 64 100
10 9 13 117 81 169
11 13 12 156 169 144
12 7 9 63 49 81
N=12 ∑ 𝑋 = 147 ∑ 𝑌 = 166 ∑ 𝑋𝑌 = 2144 ∑ 𝑋 2 = 1925 ∑ 𝑌 2 = 2412

Statistics and Probability P a g e 7 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
4th step: Compute.
a. Identify all the given Given:
𝑛 = 12 ∑ 𝑋 = 147 ∑ 𝑌 = 166

∑ 𝑋𝑌 = 2144 ∑ 𝑋 2 = 1925 ∑ 𝑌 2 = 2412

b. Use the formula 𝑛 ∑ 𝑋𝑌 − (∑ 𝑋)( ∑ 𝑌)


𝑟=
√[𝑛 ∑ 𝑋 2 − (∑ 𝑋)2 ][𝑛 ∑ 𝑌 2 − (∑ 𝑌)2 ]
c. Substitute the given and (12)(2144) − (147)(166)
𝑟=
simplify √[(12)(1925) − (147)2 ][(12)(2412) − (166)2 ]
25728 − 24402
𝑟=
√[23100 − 21609][28944 − 27556]
1326 1326
𝑟= =
√[1491][1388] √2069508]
𝑟 = 0.92
d. State the conclusion Based on the quantitative interpretation of the degree of
the linear relationship, 0.92 is found on the interval 0.91
to 0.99. It indicates that there is a very high positive
correlation between the scores in mathematics and
science. Since the result is positive correlation, the
score in math increases, the score in science also
increases or vice versa.
or we can state that:
The r value implies a very high positive correlation
between the test scores in mathematics and science.
This denotes that as the score in mathematics increases,
the score in science also increases.

Statistics and Probability P a g e 8 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Next, we test the significance of 𝑟

A. State the null and alternative Ho: There is no significant relationship between the
hypothesis test scores in Mathematics and Science.
Ha: There is a significant relationship between the
test scores in Mathematics and Science.

B. Determine the level of Level of significance: α=0.05


significance and the degree of
Degree of freedom:𝑑𝑓 = 𝑛 − 2 = 12 − 2 = 10
freedom

C. Locate for the critical values on Using α=0.05 and 𝑑𝑓 = 10, employing two-tailed
the t-Distribution table, two-tailed test the tabular value is 2.228.
Since it’s a two-tailed test, the rejection region is on
the left of -2.228 and right of 2.228.

D. Compute for the value of the t-statistics.

Identify the given n=12 𝑟 =0.92


Use the formula
𝑛−2
𝑡 = (𝑟)(√
1 − 𝑟2

Substitute and simplify


12 − 2
𝑡 = (0.92) (√ )
1 − (0.92)2

10
𝑡 = (0.92) (√ ) = (0.92)(√65.10
0.1536

𝑡 = 7.42
E. Statistical decision. Since it’s a two-tailed test, the rejection region is on
the left of -2.228 and right of 2.228.
The computed value is found on the rejection
region (right of 2.228) because 7.42 > 2.228.
Therefore, we reject the null hypothesis (Ho)

F. State the Conclusion. There is a significant relationship between the test


scores in Mathematics and Science of the 12 senior
high students.

Statistics and Probability P a g e 9 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
2. An Education Researcher wishes to determine if there is a significant relationship between
the reading comprehension and the vocabulary test among the 15 Grade 11 students of St.
Francis. Use 0.01 as the level of significance.
The scores of the13 students on the test are shown below.

Reading Comprehension Vocabulary Test


Student
(X) (Y)
1 3 11
2 7 12
3 8 19
4 9 5
5 8 17
6 9 3
7 7 15
8 10 2
9 9 15
10 5 8
11 7 12
12 6 4
13 5 10
14 8 13
15 5 5

Since we already know the steps in identifying all the given to solve for r, we can just complete
the table.

Reading Comprehension
Student Vocabulary Test (Y)
(X) 𝑋𝑌 𝑋2 𝑌2
1 3 11
2 7 12
3 8 19
4 9 5
5 8 17
6 9 3
7 7 15
8 10 2
9 9 15
10 5 8
11 7 12
12 6 4
13 5 10
14 8 13

Statistics and Probability P a g e 10 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
15 5 5
N=12 ∑𝑋 = ∑𝑌 = ∑ 𝑋𝑌 = ∑ 𝑋2 = ∑ 𝑌2 =

Complete the table:

Student Reading Vocabulary


Comprehension Test 𝑋𝑌 𝑋2 𝑌2
(X) (Y)
1 3 11 33 9 121
2 7 12 84 49 144
3 8 19 152 64 361
4 9 5 45 81 25
5 8 17 136 64 289
6 9 3 27 81 9
7 7 15 105 49 225
8 10 2 20 100 4
9 9 15 135 81 225
10 5 8 40 25 64
11 7 12 84 49 144
12 6 4 24 36 16
13 5 10 50 25 100
14 8 13 104 64 169
15 5 5 25 25 25
∑ 𝑋 = 106 ∑ 𝑌 = 151 ∑ 𝑋𝑌 = 1064 ∑ 𝑋2 = 802 ∑ 𝑌 2 = 1921
N=15

Compute for r

a. Identify all the given Given:


𝑛 = 15 ∑ 𝑋 = 106 ∑ 𝑌 = 151

∑ 𝑋𝑌 = 1064 ∑ 𝑋 2 = 802 ∑ 𝑌 2 = 1921

b. Use the formula 𝑛 ∑ 𝑋𝑌 − (∑ 𝑋)( ∑ 𝑌)


𝑟=
√[𝑛 ∑ 𝑋 2 − (∑ 𝑋)2 ][𝑛 ∑ 𝑌 2 − (∑ 𝑌)2 ]

c. Substitute the given and (15)(1064) − (106)(151)


𝑟=
√[(15)(802) − (106)2 ][(15)(1921) − (151)2 ]

Statistics and Probability P a g e 11 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
simplify 15960 − 16006
𝑟=
√[12030 − 11236][28815 − 22801]
−46 −46
𝑟= =
√[794][6014] √4775116

𝑟 = −0.02
d. State the conclusion Based on the quantitative interpretation of the degree of
the linear relationship, -0.02 is on the interval
-0.01 to -0.30 therefore it indicates that there is a
negligible negative correlation the reading
comprehension and the vocabulary of the 15 Grade 11
students of IV-Del Pilar.

Or simply we can state that:


The r value implies negligible negative correlation the
reading comprehension test and the vocabulary of the
15 Grade 11 students of St. Francis. This states that if the
reading comprehension increases the vocabulary
decreases and vice versa.

Now, we Test the significance of r for example # 2.

A. State the null and alternative Ho: There is no significant relationship between the test
hypothesis scores in reading comprehension and the vocabulary of
the 15 Grade 11 students of St. Francis
Ha: There is a significant relationship between the test
scores in reading comprehension and the vocabulary of
the 15 Grade 11 students of St. Francis

B. Determine the level of • Level of significance: α=0.01


significance and the degree of • Degree of freedom
freedom 𝑑𝑓 = 𝑛 − 2 = 15 − 2 = 13
C. Locate for the critical values of t Using α=0.01 and 𝑑𝑓 = 13, employing two-tailed test
on the t-Distribution table, two- the tabular value is 2.16
tailed

D. Compute for the value of the t-statistics.


Identify the given n=15 r=-0.02

Statistics and Probability P a g e 12 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Use the formula
𝑛−2
𝑡 = 𝑟 (√ )
1 − 𝑟2

Substitute and simplify


15 − 2
𝑡 = (−0.02) (√ )
1 − (−0.02)2

13
𝑡 = (−0.02) (√ ) = (−0.02)(3.61)
0.996

𝑡 = −0.072
E. Make a decision. Since it’s a two-tailed test, the rejection region is on the
left of -2.16 and right of 2.16.
The computed value is found on the nonrejection region
(between -2.16 and 2.16). Therefore, we do not reject
the null hypothesis (Ho)

F. State the Conclusion. There is no significant relationship between the test


scores in reading comprehension and the vocabulary of
the 15 Grade 11 students of St. Francis

3. A nutritionist wants to determine if there is a relationship between the number of hours a


person exercise and the amount of milk (in ounces) each person consumes per week.
Use 0.05 level of significance.

Respondent Number of Hours Amount of milk


No. (x) (y)
1 3 48
2 0 8
3 2 32
4 5 64
5 8 10
6 5 32
7 10 56
8 2 72
9 1 48
Since we already know the steps in identifying all the given to solve for r, we can just complete
the table.

Statistics and Probability P a g e 13 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Respondent Number of Amount of
𝑋𝑌 𝑋2 𝑌2
No. Hours (X) Milk (Y)
1 3 48 144 9 2304
2 0 8 0 0 64
3 2 32 64 4 1024
4 5 64 320 25 4096
5 8 10 80 64 100
6 5 32 160 25 1024
7 10 56 560 100 3136
8 2 72 144 4 5184
9 1 48 48 1 2304
N=9 ∑ 𝑋 = 36 ∑ 𝑌 = 370 ∑ 𝑋𝑌 = 1520 ∑ 𝑋 2 = 232 ∑ 𝑌 2 = 19236

Compute for r
a. Identify all the given Given:
𝑛=9 ∑ 𝑋 = 36 ∑ 𝑌 = 370

∑ 𝑋𝑌 = 1520 ∑ 𝑋 2 = 232 ∑ 𝑌 2 = 19236

b. Use the formula 𝑛 ∑ 𝑋𝑌 − (∑ 𝑋)( ∑ 𝑌)


𝑟=
√[𝑛 ∑ 𝑋 2 − (∑ 𝑋)2 ][𝑛 ∑ 𝑌 2 − (∑ 𝑌)2 ]
c. substitute the given and (9)(1520) − (36)(370)
𝑟=
simplify √[(9)(232) − (36)2 ][(9)(19236) − (370)2 ]
13680 − 13320
𝑟=
√[2088 − 1296][173124 − 136900]
360 360
𝑟= =
√[792][36224] √28689408

𝑟 = 0.067
d. State the conclusion Based on the quantitative interpretation of the degree of
the linear relationship, 0.067 is between - 0.01 to - 0.30
therefore it indicates that there is a negligible positive
correlation between the number of hours a person
exercise and the amount of milk (in ounces) each person
consumes per week. Since the result is positive

Statistics and Probability P a g e 14 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
correlation, it indicates that as the number of hours of
exercise increases, the amount of milk a person
consumes also increases, and vice versa.
or we can state that:
The r value implies negligible positive correlation
between the number of hours of exercise and the
amount of milk (in ounces) each person consumes per
week. This indicates that as the number of hours a
person exercise increases, the amount of milk a person
consumes also increase, and vice versa.
Now, test the significance of r for example # 3.

A. State the null and alternative Ho: There is no significant relationship between the
hypothesis number of hours of exercise and the amount of milk
(in ounces) consumed per week.
Ha: There is a significant relationship between the
number of hours of exercise and the amount of milk
(in ounces) consumed per week.

B. Determine the level of • Level of significance: α=0.05


significance and the degree of • Degree of freedom
freedom 𝑑𝑓 = 𝑛 − 2 = 9 − 2 = 7
C. Locate for the critical values of t Using α=0.05 and 𝑑𝑓 = 7, employing two-tailed
on the T-Distribution table, two- test the tabular value is 2.365
tailed

D. Compute for the value of the t-statistics.

➢ Identify the given n=9 r=-0.067

➢ Use the formula


𝑛−2
𝑡 = (𝑟) (√ )
1 − 𝑟2

➢ Substitute and simplify


9−2
𝑡 = (0.067) (√ )
1 − (0.067)2

7
𝑡 = (0.067)√
0.995

Statistics and Probability P a g e 15 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
𝑡 = 0.178
E. Make a decision. Since it’s a two-tailed test, the rejection region is on
the left of -2.365 and right of 2.365.
The computed value is found on the nonrejection
region (between -2.365 and 2.365). Therefore, we
do not reject the null hypothesis (Ho)

F. State the Conclusion. There is no significant relationship between the


number of hours of exercise and the amount of milk
(in ounces) each person consumes per week.

SIMPLE REGRESSION ANALYSIS

In studying relationships between two variables, collect the data and then construct a scatter
plot. The purpose of the scatter plot, as indicated previously, is to determine the nature of the
relationship. The possibilities include a positive linear relationship, a negative linear
relationship, a curvilinear relationship, or no discernible relationship. After the scatter plot is
drawn, the next steps are to compute the value of the correlation coefficient and to test the
significance of the relationship. If the value of the correlation coefficient is significant, the next
step is to determine the equation of the regression line, which is the data’s line of best fit.
(Note: Determining the regression line when r is not significant and then making predictions
using the regression line are meaningless.) The purpose of the regression line is to enable the
researcher to see the trend and make predictions on the basis of the data.

Line of Best Fit

Figure 1 shows a scatter plot for the data of two variables. It shows that several lines can be
drawn on the graph near the points. Given a scatter plot, you must be able to draw the line of
best fit. Best fit means that the sum of the squares of the vertical distances from each point to
the line is at a minimum. The reason you need a line of best fit is that the values of y will be
predicted from the values of x; hence, the closer the points are to the line, the better the fit and
the prediction will be. See Figure 1. When r is positive, the line slopes upward and to the right.
When r is negative, the line slopes downward from left to right.

Figure 1

Statistics and Probability P a g e 16 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Determination of the Regression Line Equation

In Algebra, the equation of a line is usually given as 𝑦 = 𝑚𝑥 + 𝑏. In Statistics, the equation of


the regression line is 𝑦 ′ = 𝑎 + 𝑏𝑥 where a is the y-intercept and b is the slope of the line.

FORMULA FOR THE REGRESSION LINE

𝒚′ = 𝒂 + 𝒃𝒙

(∑ 𝒚)(∑ 𝒙𝟐 ) − (∑ 𝒙)(∑ 𝒙𝒚) 𝑛(∑ 𝑥𝑦) − (∑ 𝑥)(∑ 𝑦)


𝒂= 𝟐
𝑏= 2
𝒏(∑ 𝒙𝟐 ) − (∑ 𝒙) 𝑛(∑ 𝑥 2 ) − (∑ 𝑥)

Note: Round the values of A and B to three decimal places.

Let’s use the data on the examples above.

Example1: A statistic teacher recorded the Math and Science test scores of twelve senior high
students from one section. Find the equation of the regression.

Student Math Scores Science Scores


XY X2 Y2
Number (X) (Y)
1 18 20 360 324 400
2 16 18 288 256 324
3 11 12 132 121 144
4 15 17 255 255 289
5 15 15 225 255 255
6 11 14 154 121 196
7 11 12 132 121 144
8 13 14 182 169 196
9 8 10 80 64 100

Statistics and Probability P a g e 17 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
10 9 13 117 81 169
11 13 12 156 169 144
12 7 9 63 94 81
N=12 ∑ 𝑋 = 147 ∑ 𝑌 = 166 ∑ 𝑋𝑌 = 2144 ∑ 𝑋 2 = 1925 ∑ 𝑌 2 = 2412

Solution:
a. Given n=6
∑ 𝑋 = 147 , ∑ 𝑌 = 166, ∑ 𝑋𝑌 = 2144 , ∑ 𝑋 2 = 1925
b. Compute for a (∑ 𝑦)(∑ 𝑥 2 ) − (∑ 𝑥)(∑ 𝑥𝑦)
𝑎= 2
𝑛(∑ 𝑥 2 ) − (∑ 𝑥)

(166)(1925) − (147)(2144)
𝑎=
6(1925) − (147)2

319550 − 315168 4382


𝑎= =
11550 − 21609 −10059

𝒂 = −𝟎. 𝟒𝟑𝟒
c. Compute for b 𝑛(∑ 𝑥𝑦) − (∑ 𝑥)(∑ 𝑦)
𝑏= 2
𝑛(∑ 𝑥 2 ) − (∑ 𝑥)

6(2144) − (147)(166)
𝑏=
6(1925) − (1472 )

12864 − 24402 −11538


𝑏= =
11550 − 21609 −10059

𝒃 = 𝟏. 𝟏𝟒𝟕
d. Substitute values of a and b 𝒚′ = 𝑎 + 𝑏𝒙
to the formula 𝒚′ = −𝟎. 𝟒𝟑𝟒 + 𝟏. 𝟏𝟒𝟕𝒙
Hence, the equation of the regression line 𝑦 ′ = 𝑎 + 𝑏𝑥 is 𝒚′ = −𝟎. 𝟒𝟑𝟒 + 𝟏. 𝟏𝟒𝟕𝒙

GENERALIZATION:
It is important to note that even if the correlation between two variables is high, it does not
necessarily mean causation. There are other possibilities, such as lurking variables or just a

Statistics and Probability P a g e 18 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
coincidental relationship. Thus, when the null hypothesis is rejected, the researcher must
consider all possibilities and select the appropriate one as determined by the study. Remember,
correlation does not necessarily imply causation. And that no regression should be done when r
is not significant.

Statistics and Probability P a g e 19 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Final Term – Second Semester Week 8: May 17-20, 2021

I. INTRODUCTION

Hello, Louisian GEM! Welcome to the final week of our Correspondence Learning! Brace
yourself as we are about to reach the finish line!

This week, you shall be given one, last lesson to study. Attached to this module is the
Weekly Study and Assessment Guide

DATE TOPIC ACTIVITIES OR TASKS


Chi-square test of Goodness-of-Fit Read the lesson.
May 17 – 20, 2021
Chi-square test of Independence

For this week of the final term, the following shall be your guide for the different lessons
and tasks that you need to accomplish. Be patient, read it carefully before proceeding to the
tasks expected of you. GOOD LUCK!

Content Chi-square Test


- One- Variable Chi-Square Test
- Two- Variable Chi-Square Test
Learning Competencies The learner should be able to:
• explain how the test statistic is computed for a chi-square
test;
• calculate the degrees of freedom for the chi-square
goodness-of-fit test and locate critical values in the chi-
square table;
• compute the chi-square goodness-of-fit test and interpret
the results;
• identify the assumption and the restriction of expected
frequency size for the chi-square test;
• calculate the degrees of freedom for the chi-square test for
independence and locate critical values in the chi-square
table;
• compute the chi-square test for independence and
interpret the results; and
• relate the steps in hypothesis testing in making a sound
judgment and better decision making in real life situations.
Activities Read the lesson
Essential Questions When do you know that you made the best and right decision?
Value Statement You cannot make progress without making decisions.
- Jim Rohn
One finds the truth by making a hypothesis and comparing

Statistics and Probability P a g e 20 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
observations with the hypotheses.
- David Douglass
References Textbooks:

Altares, P. et, al. (2012). Elementary Statistics with Computer


Applications. Rex Book store. Quezon City, Philippines.

Icutan, S. L. et, al. (2012). Statistics with Probability a


Comprehensive Approach. Jimczyville Publications. Malabon City,
Philippines

Cabrera, R. et, al. (2014) Fundamental of Statistics. Grandbooks


Publishing. Manila, Philippines

II. LEARNING CONTENT

We have already covered the different parametric tests in the past weeks of our
correspondence learning. We are now going to consider one of the most widely used non-
parametric test in the world of statistics which is the Chi-Square Test.

Most discussions of statistical procedures have been focused on the analysis of “quantitative”
data, but not all the variables in research are quantitative. Instead, they are “qualitative” in
nature.

Recall that: A qualitative data is one where the possible measurements are not measurements
but quantities or frequencies of things that occur in categories. Religious affiliation, gender, or
whichever brand of dishwashing detergent you prefer are examples.

A very useful test of this type is the Chi-Square test. It is used with ordinal data in the form of
frequencies or proportions. It uses counts or frequencies as data rather than means and
standard deviations. It requires the use of numerical values.

Concepts of Contingency Table

The 2 x 2 contingency table is one of the most common ways to summarize two qualitative
variables. To illustrate it, consider a sample of COVID-19 cases from Community A that was
classified according to symptoms and their sex.

Symptoms
Sex Total
Asymptomatic Symptomatic
Male 17 14 31
Female 10 9 19
Total 27 23 50

Statistics and Probability P a g e 21 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
The generic 2x2 contingency table includes row marginal total, column marginal total, and
grand total.
The chi-square statistic is used to compare the observed frequency of some observation (such
as frequency of buying different brands of laptops) with an expected frequency (such as buying
equal numbers of each brand of laptop. The comparison of observed and expected frequencies
is used to calculate the value of the chi-square statistic, which in turn can be compared with the
distribution of chi square to make inference about a statistical problem.

Chi square is computed using the formula:

(𝑂 − 𝐸 ) 2 where:
𝑥2 = 𝛴 𝑂 = 𝑜𝑏𝑠𝑒𝑟𝑣𝑒𝑑 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦
𝐸
𝐸 = 𝑒𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦

The degrees of freedom for the one-variable where:


chi-square statistic is:
𝑑𝑓 = 𝑐 − 1 𝑐 = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑐𝑎𝑡𝑒𝑔𝑜𝑟𝑖𝑒𝑠 𝑜𝑟 𝑙𝑒𝑣𝑒𝑙𝑠

As for the steps in hypothesis testing involving chi-square test, we follow the same steps in
hypothesis testing just like our previous lessons for the past weeks.
We recall the steps in hypothesis testing:
a. State the null and alternative hypotheses.
b. Select the level of significance.
c. Identify the appropriate test-statistic and determine whether it is one-tailed or two-tailed.
*note that chi-square test is always a right-tailed test since the critical values in the chi-square distribution are
all positive.
d. Determine the critical value and the rejection region.
*The chi-square distribution is being used. (The chi-square distribution table is found on the next page).
e. Compute the test statistic (use chi-square test formula)
f. State the decision rule
Reject the null hypothesis (Ho) if the computed value of the statistic is greater than the critical value.
Otherwise, do no reject the null hypothesis.
g. Conclusion
- Based on the test results, the researcher will reach a conclusion about the population under study.

Statistics and Probability P a g e 22 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Statistics and Probability P a g e 23 | 33
This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Let’s get started!

Chi-Square Test of Goodness-of-Fit

The chi-square test of goodness-of-fitt is used if the researcher is interested in determining the
number of objects or responses which fall in different categories for a single qualitative variable
and wish to know if sample under analysis was drawn from a proportion with some specified
distribution or not.

Let’s analyze the following examples.

Example 1:
The data for 150 students is recorded in the table below (the observed frequencies). We have
also indicated the expected frequency for each category. Since there are 150 measures or
observations and there are three categories (Dell, Asus, and Acer) we would indicate the
expected frequency for each category (same for all level) to be 150/3 or 50. Use 𝛼 = 0.05
significance level.

Computer Observed frequency Expected Frequency


Dell 27 50
Asus 59 50
Acer 64 50

We now have the information we need to complete the steps in testing our statistical
hypotheses.

a. State the null and alternative Ho: There is no significant difference between the
hypotheses observed and the expected frequencies (we can
also say that the respondents show no preference)

Ha: There is a significant difference between


observed and expected frequencies (we can also
say that the respondents show preference)

b. Select the level of significance 𝛼 = 0.05

c. Identify the appropriate test- Chi- Square (always right tailed)


statistic and determine whether
it is one-tailed or two-tailed

𝛼 = 0.05, 𝑑𝑓 = 𝑐 − 1 = 3 − 1 = 2
d. Determine the critical value and
(3 categories → Dell, Asus, Acer)

Statistics and Probability P a g e 24 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
the rejection region Critical value: 5.991
*check the chi-square distribution table

e. Compute the test statistic


Frequency with which students select Computer
Brand
Computer Observed Expected (𝑂 − 𝐸)2
frequency Frequency 𝐸

Dell 27 50 10.58
Asus 59 50 1.62
Acer 64 50 3.92

2
(𝑂 − 𝐸 ) 2
𝑥 = 𝛴
𝐸
( 27 − 50)2 (59 − 50)2 (64 − 50)2
𝑥2 = + +
50 50 50
𝑥 2 = 10.58 + 1.62 + 3.92 = 16.12
f. State the statistical decision. Since computed (16.12) is greater than the critical
value (5.991), we reject the null hypothesis (Ho).
g. Conclusion There is a significant difference between the
observed and expected frequencies. We can also
say that the respondents show preference when it
comes to the brand of laptop they are going to buy.

Example 2:
There are three gates at the certain university. The building maintenance supervisor would like
to know if the gates are equally utilized. As an experiment, 600 are observed as they enter the
school. The number of students using each gate is reported below. At 0.01 significance level, can
we conclude that there is a difference in the use of the three gates?
Gate Number of students
Gate 1 245
Gate 2 205
Gate 3 150
Total 600

Statistics and Probability P a g e 25 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Because there are 600 students in the sample, we expect that 200 students fall in each of the
three categories.

Gate Observed Frequency Expected Frequency


Gate 1 245 200
Gate 2 205 200
Gate 3 150 200
Total 600 600

We now proceed to the steps in hypothesis testing.


Solution:
a. State the null and alternative Ho: There is no significant difference between the
hypotheses observed and expected frequencies (the gates are
equally utilized)

Ha: There is a significant difference between the


observed and expected frequencies (the gates are
not equally utilized)

b. Select the level of significance 𝛼 = 0.01

c. Identify the appropriate test- Chi- Square test, right-tailed


statistic and determine whether
it is one-tailed or two-tailed

d. Determine the critical value and Since,


the rejection region 𝛼 = 0.01
𝑑𝑓 = 𝑐 − 1 (3 categories → gate1, gate2, gate3)
𝑑𝑓 = 3 − 1 = 2

Critical value = 9.210


*check the chi-square distribution table
e. Compute the test statistic
Frequency of use among the Gates in a
University
Computer Observed Expected (𝑂 − 𝐸)2
frequency Frequency 𝐸

Gate 1 245 200 10.125


Gate 2 205 200 0.125
Gate 3 150 200 12.5

Statistics and Probability P a g e 26 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
(𝑂 − 𝐸 ) 2
𝑥2 = 𝛴
𝐸
(245 − 200)2 (205 − 200)2 (150 − 200)2
= + +
200 200 200
2
𝑥 = 10.125 + 0.125 + 12.5 = 22.75
f. State the statistical decision. Since the computed value (22.75) is greater than
the critical value (9.210), we reject the null
hypothesis (Ho).
There is a significant difference between the
g. Conclusion observed and expected frequencies. The three
gates are not equally utilized.

STATrivia!
Did you know that chi-square goodness-of-fit test was developed by Karl Pearson? The purpose
of this test is to determine how well an observed set of data fits an expected data.

Chi-square Test of Independence

In this test, two variables are involved and one variable is tested for independence to the other
variable.
Now let us consider the case of the two-variable chi-square test, also known as the Test of
Independence.
Example 1:
For example, we may wish to know if there is an association between the employment status
and the type of school where the employees graduated from. Use 0.05 level of significance.

Type of School Graduated From Total


State
Local
Employment Colleges or Private Private Non-
Government
Universities Sectarian Sectarian
Unit Colleges
(SCUs)
Hired 46 11 23 25 105
Not Hired 21 3 9 12 45
Total 67 14 32 37 150

The formula for chi-square is the same as before only that the degrees of freedom of a two-
variable chi square statistic is:

Statistics and Probability P a g e 27 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
where:
𝐷𝑒𝑔𝑟𝑒𝑒𝑠 𝑜𝑓 𝐹𝑟𝑒𝑒𝑑𝑜𝑚
𝐶 = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑐𝑜𝑙𝑢𝑚𝑛𝑠 𝑜𝑟 𝑙𝑒𝑣𝑒𝑙𝑠 𝑜𝑓 𝑡ℎ𝑒 𝑓𝑖𝑟𝑠𝑡 𝑣𝑎𝑟𝑖𝑎𝑏𝑙𝑒
𝑑𝑓 = (𝐶 − 1)(𝑅 − 1)
𝑅 = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑟𝑜𝑤𝑠 𝑜𝑟 𝑙𝑒𝑣𝑒𝑙𝑠 𝑜𝑓 𝑡ℎ𝑒 𝑠𝑒𝑐𝑜𝑛𝑑 𝑣𝑎𝑟𝑖𝑎𝑏𝑙𝑒

To solve for the expected frequency in Chi-Square Test of Independence:


(𝑟𝑜𝑤 𝑡𝑜𝑡𝑎𝑙)(𝑐𝑜𝑙𝑢𝑚𝑛 𝑡𝑜𝑡𝑎𝑙)
𝐸𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦 =
𝐺𝑟𝑎𝑛𝑑 𝑡𝑜𝑡𝑎𝑙

From the given, the grand total is 150

(105)(67) (105)(14) (32)(105) (37)(105)


𝐸1 = = 46.9 𝐸3 = = 9.8 𝐸5 = = 22.4 𝐸7 = = 25.9
150 150 150 150

(45)(67) (14)(45) (32)(45) (37)(45)


𝐸2 = = 20.1 𝐸4 = = 4.2 𝐸6 = = 9.6 𝐸8 = = 11.1
150 150 150 150

Let’s input the expected frequencies on the contingency table.

Type of School Graduated From


State Colleges or Local Government Private Non-
Employment Private Sectarian Total
Universities (SCUs) Unit Colleges Sectarian
Observed Expected Observed Expected Observed Expected Observed Expected
Hired 46 46.9 11 9.8 23 22.4 25 25.9 105
Not Hired 21 20.1 3 4.2 9 9.6 12 11.1 45
Total 67 14 32 37 150

All the needed data are already complete, we’re now going to test our hypothesis.
Solution:

a. State the null and alternative Ho: There is no association between the employment status and the type of
hypotheses school.

Ha: There is association between the employment status and the type of
school.

b. Select the level of significance 𝛼 = 0.05

c. Identify the appropriate test- Chi- Square (always right tailed)


statistic and determine whether
it is one-tailed or two-tailed

d. Determine the critical value Since,


𝛼 = 0.05
𝑑𝑓 = (C − 1)(R − 1)

Statistics and Probability P a g e 28 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
and the rejection region C = SCUs, LGUC, Private Sectarian, Non − Private Sectarian ∶ 4
R = Hired, Not Hired: 2
𝑑𝑓 = (4 − 1)(2 − 1) = (3)(1) = 3
Critical value = 7.815

e. Compute the test statistic


STATE COLLEGES OR LOCAL PRIVATE SECTARIAN PRIVATE NON- TOTAL
UNIVERSITIES (SCUS) GOVERNMENT SECTARIAN

Employment

0 E (𝑂 − 𝐸)2 0 E (𝑂 − 𝐸)2 0 E (𝑂 − 𝐸)2 0 E (𝑂 − 𝐸)2


𝐸 𝐸 𝐸 𝐸
Hired 46 46.9 0.02 11 9.8 0.15 23 22.4 0.02 25 25.9 0.03 105

Not Hired 21 20.1 0.04 3 4.2 0.34 9 9.6 0.04 12 11.1 0.07 45

Total 67 14 32 37 150

(𝑂 − 𝐸 ) 2
𝑥2 = 𝛴 = 0.02 + 0.04 + 0.15 + 0.34 + 0.02 + 0.04 + 0.03 + 0.07 = 0.71
𝐸

Since the computed value (0.71) is less than the critical value (7.815), do not
f. State the decision rule reject the null hypothesis (Ho).

g. Conclusion There is no association between the employment status and the type of
school they graduated from.

Statistics and Probability P a g e 29 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Example 4:
A survey in which independent random samples of 80 single persons, 120 married persons, and
100 person who are widowed were asked whether “friends and social life”, “job or primary
activity”, or “health and physical condition” contribute most to their general happiness. The
results are shown below. Use 𝛼 = 0.05 significance level.

Social Status
Factors Affecting General Happiness TOTAL
SINGLE MARRIED WIDOWED
Friends and Social Life 41 49 42 132
Job or Primary Activity 27 50 33 110
Health and Physical Condition 12 21 25 58
TOTAL 80 120 100 300

(𝑟𝑜𝑤 𝑡𝑜𝑡𝑎𝑙)(𝑐𝑜𝑙𝑢𝑚𝑛 𝑡𝑜𝑡𝑎𝑙)


𝐸𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦 =
𝐺𝑟𝑎𝑛𝑑 𝑡𝑜𝑡𝑎𝑙
(132)(80) (132)(120) (132)(100)
𝐸1 = = 35.2 𝐸4 = = 52.8 𝐸7 = = 44
300 300 300

(110)(80) (110)(120) (110)(100)


𝐸2 = = 29.33 𝐸5 = = 44 𝐸8 = = 36.67
300 300 150
(58)(80) (58)(120) (58)(100)
𝐸3 = = 15.47 𝐸6 = = 23.2 𝐸9 = = 19.33
300 300 150

We now have the information we need to complete the steps in testing our statistical
hypotheses for our research problem.
Solution:
Ho: There is no association between the social status and the factors
a. State the null and alternative affecting the general happiness.
hypotheses Ha: There is an association between the social status and the factors
affecting the general happiness.

𝛼 = 0.05
b. Select the level of significance

c. Identify the appropriate test- Chi- Square (always right tailed)


statistic and determine whether it
is one-tailed or two-tailed

Statistics and Probability P a g e 30 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
c. Determine the critical value and Since,
the rejection region 𝛼 = 0.01
𝑑𝑓 = (𝐶 − 1)(𝑅 − 1)
𝐶 = 𝑠𝑖𝑛𝑔𝑙𝑒, 𝑚𝑎𝑟𝑟𝑖𝑒𝑑, 𝑤𝑖𝑑𝑜𝑤𝑒𝑑 (3)
𝑅 = 𝐹𝑟𝑖𝑒𝑛𝑑𝑠 𝑎𝑛𝑑 𝑆𝑜𝑐𝑖𝑎𝑙 𝐿𝑖𝑓𝑒 , 𝐽𝑜𝑏 𝑜𝑟 𝑃𝑟𝑖𝑚𝑎𝑟𝑦 𝐴𝑐𝑡𝑖𝑣𝑖𝑡𝑦,
𝐻𝑒𝑎𝑙𝑡ℎ 𝑎𝑛𝑑 𝑃ℎ𝑦𝑠𝑖𝑐𝑎𝑙 𝐶𝑜𝑛𝑑𝑖𝑡𝑖𝑜𝑛 (3)
𝑑𝑓 = (3 − 1)(3 − 1) = (2)(2)
𝑑𝑓 = 4
Critical value = 9.488

d. Compute the test statistic.


SINGLE MARRIED WIDOWED
( 𝑂 − 𝐸 )2 (𝑂 − 𝐸 ) 2 (𝑂 − 𝐸 )2 TOTAL
0 E 0 E 0 E
𝐸 𝐸 𝐸
Friends and
41 35.20 0.96 49 52.80 0.27 42 44.00 0.09 132
Social Life
Job or
Primary 27 29.33 0.19 50 44.00 0.82 33 36.67 0.37 110
Activity
Health and
Physical 12 15.47 0.78 21 23.20 0.21 25 19.33 1.66 58
Condition
TOTAL 80 120 100 300
(𝑂 − 𝐸 ) 2
𝑥2 = 𝛴 = 0.96 + 0.19 + 0.78 + 0.27 + 0.82 + 0.21 + 0.09 + 0.37 + 1.66 = 5.35
𝐸

e. State the decision rule Since the computed value of 𝑥 2 (5.35) is less than the critical value (9.488),
do not reject Ho.
There is no association between the social status and the factors affecting
f. Conclusion
the general happiness.

Statistics and Probability P a g e 31 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
STATrivia!

When there is a close agreement between the value of O and E, the chi-square is small and the null
hypothesis is not rejected. When there are large differences between O and E, the chi-square is
large and the null hypothesis is rejected.

Before we finally close this final chapter of our correspondence learning, we should always bear
in our mind the essence of why we need to learn and study the different hypothesis testing that
we have had this whole semester. All of these things will train our mind in critical thinking
which will lead us to a better and wiser decision making. With these, we are not just thought to
solve problems but to infer better and deeper understanding of the situations that are
presented in front of us. We are able to unravel facts that we thought impossible. We are able
correct misconceptions that we always commit. Also, we discover things that yet to be
discovered by ourselves. Who knows that this small discovery of ours will change our life
forever? Most especially, we are able to weigh things out and decide what to reject and do not
accept in our life

Generalization:

Important points of this Week’s Correspondence Learning:

Chi-square test is used when your data is nominal or ordinal which falls under qualitative. It is
used to compare the differences of the sample frequencies with the expected frequencies. It
has two applications namely: test of goodness-of-fit and test of association or independence.

Chi square (x2) is computed using the formula:

where:
2
(𝑂 − 𝐸 ) 2 𝑂 = 𝑜𝑏𝑠𝑒𝑟𝑣𝑒𝑑 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦
𝑥 =𝛴
𝐸 𝐸 = 𝑒𝑥𝑝𝑒𝑐𝑡𝑒𝑑 𝑓𝑟𝑒𝑞𝑢𝑒𝑛𝑐𝑦

The degrees of freedom for the one-variable where:


chi-square statistic is:
𝑐 = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑐𝑎𝑡𝑒𝑔𝑜𝑟𝑖𝑒𝑠 𝑜𝑟 𝑙𝑒𝑣𝑒𝑙𝑠
𝑑𝑓 = 𝑐 − 1
The degrees of freedom for the two-variable where:
chi-square statistic is: 𝐶 = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑐𝑜𝑙𝑢𝑚𝑛𝑠 𝑜𝑟 𝑙𝑒𝑣𝑒𝑙𝑠 𝑜𝑓 𝑡ℎ𝑒 𝑓𝑖𝑟𝑠𝑡 𝑣𝑎𝑟𝑖𝑎𝑏𝑙𝑒
𝑑𝑓 = (𝐶 − 1)(𝑅 − 1) 𝑅 = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑟𝑜𝑤𝑠 𝑜𝑟 𝑙𝑒𝑣𝑒𝑙𝑠 𝑜𝑓 𝑡ℎ𝑒 𝑠𝑒𝑐𝑜𝑛𝑑 𝑣𝑎𝑟𝑖𝑎𝑏𝑙𝑒

I hope you’ve had a very meaningful school year! I pray for the success of your academic
journey. May you have a meaningful and beautiful life ahead of you.

Statistics and Probability P a g e 32 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.
Sometimes it’s
the smallest decisions
that can change your
life forever.
-Keri Russel

Statistics and Probability P a g e 33 | 33


This document is a property of University of Saint Louis Tuguegarao. It must not be reproduced or transmitted in any form, in whole or in part, without expressed written permission.

You might also like