0% found this document useful (0 votes)
4 views133 pages

Module 6 - Hypothesis Testing - Updated

This document outlines the principles and procedures of hypothesis testing in research methodology, emphasizing the role of hypotheses in guiding investigations and decision-making. It covers key concepts such as null and alternative hypotheses, significance levels, types of errors, and the importance of test power. The document also details the systematic steps involved in hypothesis testing, from formulating hypotheses to interpreting results.

Uploaded by

m2210036
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views133 pages

Module 6 - Hypothesis Testing - Updated

This document outlines the principles and procedures of hypothesis testing in research methodology, emphasizing the role of hypotheses in guiding investigations and decision-making. It covers key concepts such as null and alternative hypotheses, significance levels, types of errors, and the importance of test power. The document also details the systematic steps involved in hypothesis testing, from formulating hypotheses to interpreting results.

Uploaded by

m2210036
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

RESEARCH METHODOLOGY

Module 6 – Hypothesis Testing


Reference: Research Methodology: Methods and Techniques – C.R. Kothari

[Link]
Hypothesis Testing
Hypothesis in Research
▪ A hypothesis serves as a primary instrument in research, guiding the
direction of studies and investigations

▪ Its main function is to suggest new experiments, observations, and lines


of inquiry

▪ Many experiments are conducted with the specific objective of testing


hypotheses

▪ In decision-making scenarios, hypotheses are tested using available data


to support rational conclusions and actions

2
Hypothesis Testing
Hypothesis in Research
▪ In social sciences, where population parameters are rarely known,
hypothesis testing is a commonly used strategy for generalization

▪ It helps determine whether sample data provide sufficient support for a


given assumption

▪ Hypothesis testing enables researchers to make probabilistic statements


about population parameters

▪ A hypothesis is not proven absolutely, but is accepted if it withstands


rigorous statistical testing

▪ It is therefore a tool for inference rather than absolute proof


3
Hypothesis Testing
What is Hypothesis?
▪ A hypothesis is generally understood as an assumption or proposition to
be tested or verified

▪ In research, it is a formal and structured question that the investigator


seeks to resolve

▪ A hypothesis is a proposition or set of propositions that explains a group


of phenomena

▪ It may act as a tentative explanation guiding investigation or a statement


supported by existing evidence

4
Hypothesis Testing
What is Hypothesis?
▪ A research hypothesis is often a predictive statement linking independent
and dependent variables

▪ It indicates expected relationships that can be tested scientifically

▪ Examples:
“Students receiving counselling will show greater creativity than those
who do not.”
“Automobile A performs as well as Automobile B.”

▪ A hypothesis clearly states what the researcher is looking for and must be
capable of empirical testing

5
Hypothesis Testing
Characteristics
▪ Clarity and Precision: A hypothesis should be clear, specific, and precise,
ensuring reliable interpretation of results

▪ Testability: It must be capable of being tested through observation or


experimentation. Untestable hypotheses can lead to ineffective or stalled
research efforts

▪ Relationship Between Variables: A good hypothesis should explicitly


state the relationship between variables, especially in analytical studies

▪ Specific and Limited Scope: It should be narrow in scope and well-


defined, as focused hypotheses are easier to test

6
Hypothesis Testing
Characteristics
▪ Simplicity: The hypothesis should be simple and easy to understand,
without compromising its significance

▪ Consistency with Existing Knowledge: It must align with established


facts and existing theoretical frameworks

▪ Feasibility: It should be testable within a reasonable time frame, avoiding


impractical research designs

▪ Explanatory Power: A hypothesis should explain the phenomenon it


addresses, linking theory with observed facts

7
Hypothesis Testing
Basic Concepts
Null hypothesis and alternative hypothesis

▪ The null hypothesis (H₀) represents a default assumption, typically


stating no effect, no difference, or equality

▪ The alternative hypothesis (Hₐ) represents the opposite claim, indicating


the presence of an effect or difference

▪ Example: Suppose we want to test the hypothesis that the population


mean μ is equal to the hypothesized mean μHo

8
Hypothesis Testing
Basic Concepts
Null hypothesis and alternative hypothesis

▪ If our sample results do not support this null hypothesis, we should


conclude that something else is true

▪ What we conclude rejecting the null hypothesis is known as alternative


hypothesis

▪ In other words, the set of alternatives to the null hypothesis is referred to


as the alternative hypothesis

9
Hypothesis Testing
Basic Concepts
Null hypothesis and alternative hypothesis

▪ Possible alternative hypotheses:

10
Hypothesis Testing
Basic Concepts
Null hypothesis and alternative hypothesis

▪ Hypotheses must be formulated before data collection to avoid bias

▪ Following considerations to be made when choosing null hypothesis:

▪ Alternative hypothesis is usually the one which one wishes to prove and the
null hypothesis is the one which one wishes to disprove

▪ If the rejection of a certain hypothesis when it is actually true involves great


risk, it is taken as null hypothesis because then the probability of rejecting it
when it is true is α (the level of significance) which is chosen very small

▪ Null hypothesis should always be specific hypothesis i.e., it should not state
about or approximately a certain value
11
Hypothesis Testing
Basic Concepts
Why do we proceed on the basis of null hypothesis?

On the assumption that null hypothesis is true, one can assign the
probabilities to different possible sample results, but this cannot be done if
we proceed with the alternative hypothesis. Hence the use of null hypothesis
(at times also known as statistical hypothesis) is quite frequent

12
Hypothesis Testing
Basic Concepts
Level of significance

▪ The level of significance is the maximum probability of rejecting a true


null hypothesis

▪ Commonly chosen values are 5% (0.05) or 1% (0.01)

▪ A 5% significance level means there is a 5% risk of incorrectly rejecting H₀


when it is actually true

▪ It is always fixed before conducting the test, ensuring objectivity in


decision-making

13
Hypothesis Testing
Basic Concepts
Level of significance

▪ Example: A factory claims that only 2% of its light bulbs are defective. This
claim becomes the null hypothesis:
H₀ (null hypothesis): Defect rate = 2%
H₁ (alternative hypothesis): Defect rate > 2%
A quality inspector randomly samples 200 bulbs and tests them.

The inspector decides to use a 5% level of significance (α = 0.05).


This means: They are willing to accept a 5% chance of incorrectly rejecting H₀ -
i.e., concluding that the defect rate is higher than 2% when in reality it is actually
2%.
14
Hypothesis Testing
Basic Concepts
Level of significance

▪ Even if the factory is actually correct (true defect rate = 2%), there is still a
small chance (≤ 5%) that random sampling could produce an unusually high
number of defects like this. If that rare event happens, the inspector would
still reject H₀.

▪ So, the 5% significance level means:


“I accept that in about 5 out of 100 similar experiments, I might wrongly
conclude the defect rate is higher than it really is.”

15
Hypothesis Testing
Basic Concepts
Decision Rule in Hypothesis Testing

▪ A decision rule specifies criteria for accepting or rejecting H₀ based on


sample results

▪ It involves defining threshold values or conditions before testing begins

▪ Example: If testing a batch of items, we may accept H₀ if 0 or 1 defective


item is found, otherwise reject it

▪ This rule provides a systematic basis for statistical decisions.

16
Hypothesis Testing
Basic Concepts
Type I and Type II Errors

▪ Type I error (α error): Rejecting H₀ when it is actually true

▪ Type II error (β error): Accepting H₀ when it is actually false

▪ These errors represent incorrect decisions in hypothesis testing

▪ Reducing Type I error often increases the probability of Type II error,


creating a trade-off

▪ The acceptable balance depends on the cost or consequences of each


type of error

17
Hypothesis Testing
Basic Concepts
Type I and Type II Errors

18
Hypothesis Testing
Basic Concepts
Type I and Type II Errors

▪ With a fixed sample size (n), reducing Type I error (α) increases Type II
error (β)

▪ Both errors cannot be minimized simultaneously — there is a


fundamental trade-off

▪ Lowering one type of error generally leads to a rise in the other

▪ The balance between α and β depends on the real-world consequences of


each error

▪ In practical settings (business/engineering), decisions are based on the


cost of errors 19
Hypothesis Testing
Basic Concepts
Type I and Type II Errors

▪ For example, if a Type I error results in unnecessary effort, such as


reworking a batch of chemicals that was actually acceptable, the cost may
be limited to time and resources

▪ On the other hand, a Type II error in the same context could be far more
severe, such as allowing a defective chemical batch to pass, potentially
causing serious harm or danger to users

▪ In such high-risk scenarios, it is often preferable to accept a higher


probability of Type I error in order to minimize the risk of a more
dangerous Type II error 20
Hypothesis Testing
Basic Concepts
Type I and Type II Errors

▪ Consequently, decision-makers may intentionally set a higher


significance level (α) when the consequences of Type II errors are critical

▪ Overall, effective hypothesis testing requires a careful and informed


balance between Type I and Type II errors, ensuring that the chosen level
of risk aligns with the real-world implications of each type of mistake

21
Hypothesis Testing
Basic Concepts
Two-tailed and One-tailed tests

▪ A two-tailed test rejects the null hypothesis if, say, the sample mean is
significantly higher or lower than the hypothesized value of the mean of
the population

▪ Such a test is appropriate when the null hypothesis is some specified


value and the alternative hypothesis is a value not equal to the specified
value of the null hypothesis

▪ Thus, in a two-tailed test, there are two rejection regions*, one on each
tail of the curve

22
Hypothesis Testing
Basic Concepts
Two-tailed and One-tailed tests

23
Hypothesis Testing
Basic Concepts
Z (Test Statistic)

▪ Z tells you how many standard deviations your sample result is away from
the hypothesized population mean

▪ Mathematically, it is computed as:

24
Hypothesis Testing
Basic Concepts
Two-tailed and One-tailed tests

▪ If the significance level is 5 percent and the two-tailed test is to be applied,


the probability of the rejection area will be 0.05 (equally split on both tails of
the curve as 0.025) and that of the acceptance region will be 0.95 as shown in
the above curve
▪ If we take m = 100 and if our sample mean deviates significantly from 100 in
either direction, then we shall reject the null hypothesis; but if the sample
mean does not deviate significantly from m , in that case we shall accept the
null hypothesis 25
Hypothesis Testing
Basic Concepts
Two-tailed and One-tailed tests

▪ One-tailed test would be used when we are to test, say, whether the
population mean is either lower than or higher than some hypothesised
value

▪ Summary: A two-tailed test checks for differences in both directions.


For example: “Is the mean different from 100?” Here, it could be higher or
lower.
A one-tailed test checks only one direction: “Is it greater than 100?”“Is it
less than 100?”

26
Hypothesis Testing
Basic Concepts
Two-tailed and One-tailed tests

27
Hypothesis Testing
Basic Concepts
Two-tailed and One-tailed tests

28
Task
Two-Tailed Test (Material Strength)

A material is supposed to have mean strength of 500 MPa.

Sample: Mean = 480 MPa | SD = 40 MPa | n = 25

Perform the test at 5% level.

29
Task
A manufacturer claims that a temperature sensor has a mean error of 0°C
(perfect calibration).

An engineer tests n = 64 sensors and finds:


Sample mean error = +0.8°C
Known standard deviation = 2°C

Questions: Formulate H0 and Ha


Perform a two-tailed Z-test at 5% significance
Interpret the result physically
Should the sensor batch be accepted for precision applications?

30
Hypothesis Testing
Procedure
▪ Hypothesis testing involves deciding, based on collected data, whether a
hypothesis is valid

▪ The key question:→ Should we accept or reject the null hypothesis (H₀)?

▪ The procedure consists of systematic steps to choose between these two


decisions.

31
Hypothesis Testing
Procedure
1. Making a Formal Statement

▪ Clearly define: Null hypothesis (H₀) | Alternative hypothesis (Hₐ)

▪ Hypotheses must align with the research problem

▪ Examples: Bridge load capacity: H₀: μ = 10 tons | Hₐ: μ > 10 tons


Education system evaluation: H₀: μ = 80 | Hₐ: μ ≠ 80

▪ Key points:
Proper formulation is crucial
Determines the type of test:
One-tailed test → Hₐ is “greater than” or “less than”
Two-tailed test → Hₐ is “not equal to” 32
Hypothesis Testing
Procedure
2. Selecting a Significance Level

▪ Hypotheses are tested at a predetermined significance level

▪ Common choices: 5% (0.05) | 1% (0.01)

▪ Factors affecting α:
Difference between sample means
Sample size
Variability within samples
Nature of hypothesis: Directional | Non-directional
The chosen level should suit the purpose of the study

33
Hypothesis Testing
Procedure
3. Deciding the Distribution to Use

▪ Choose the appropriate sampling distribution:


Normal (Z) distribution
t-distribution

▪ Selection depends on conditions similar to estimation problems

4. Selecting a Random Sample & Computing Test Statistic

▪ Draw a random sample

▪ Compute a test statistic using sample data

▪ Use the selected distribution to obtain this value 34


Hypothesis Testing
Procedure
5. Calculating the Probability

▪ Determine the probability of obtaining the observed result (or more


extreme), assuming H₀ is true

▪ This is essentially the p-value

6. Comparing Probability with Significance Level

▪ Compare p-value with α:

▪ Decision rule: If p ≤ α → Reject H₀ (accept Hₐ)


If p > α → Do not reject (accept) H₀

▪ For two-tailed tests: Compare with α/2 35


Hypothesis Testing
Procedure
7. Understanding Errors

▪ Rejecting H₀ when true → Type I error


Risk = α (significance level)

▪ Accepting H₀ when false → Type II error


Risk cannot be precisely determined if H₀ is vague

36
Hypothesis Testing
Procedure

37
Hypothesis Testing
Power of a test
Significance Level (α) is usually fixed in advance

Primary concern then shifts to Type II Error (β):

▪ Since hypothesis tests are not perfect, there is always some possibility of
failing to reject a false null hypothesis

▪ This failure represents a Type II error, which researchers try to minimize


as much as possible

▪ Ideally, β should be very small, meaning the probability of accepting a


false null hypothesis should be low

38
Hypothesis Testing
Power of a test
Focus is often placed on (1 − β), called the Power of the Test:

▪ Instead of emphasizing β directly, statisticians often examine 1 − β, which


represents the probability of correctly rejecting a false null hypothesis

▪ This probability is called the power of the test

▪ It measures the effectiveness or sensitivity of the test in detecting a false


null hypothesis

Interpretation of power (1 − β):

▪ If 1 − β is close to 1, the test has high power, meaning it performs well in


detecting false null hypotheses
39
Hypothesis Testing
Power of a test
▪ A high-power test is desirable because it rarely overlooks real effects or
differences

▪ If 1 − β is close to 0, the test has poor power, meaning it frequently fails


to reject false null hypotheses

▪ Thus, power is a direct measure of how well a hypothesis test works

▪ The larger the power, the greater the test’s ability to detect departures
from the null hypothesis

▪ A powerful test provides stronger assurance that important effects will


not be missed
40
Hypothesis Testing
Power of a test
▪ For this reason, maximizing power is a major objective in designing
statistical tests

▪ If we compute 1 − β for every possible value of a population parameter


(for example, the true population mean μ) for which the null hypothesis is
false, and plot these probabilities, the resulting graph is called the power
curve

▪ The power curve shows how the probability of rejecting H₀ changes as the
true parameter value moves away from the hypothesized value

41
Hypothesis Testing
Power of a test
Power Function:

▪ The mathematical function that defines the power curve is called the
power function

▪ For every possible value of the population parameter, the power function
gives the probability that H₀ will be rejected

▪ The value of this function at a particular parameter point is called the


power of the test at that point

▪ As the true population parameter gets closer to the hypothesized value


under H₀, it becomes harder to distinguish the alternative from the null
hypothesis 42
Hypothesis Testing
Power of a test
Power Function:

▪ Consequently, the power of the test decreases

▪ In fact, when the true parameter exactly equals the hypothesized value,
the probability of rejecting H₀ equals the significance level α

▪ Therefore, the power curve terminates at a point whose height is α


directly above the hypothesized parameter value

Operating Characteristic Function:

▪ Closely related to the power function is another function called the


Operating Characteristic (OC) Function
43
Hypothesis Testing
Power of a test
▪ Instead of showing the probability of rejecting H₀, it shows the probability
of accepting (or not rejecting) H₀ for all possible parameter values

▪ It describes how likely the test is to retain the null hypothesis under
different conditions

▪ If the power function is denoted by H, and the operating characteristic


function by L, then: L = 1 − H

44
Hypothesis Testing
Power of a test
▪ Example:

45
Hypothesis Testing
Power of a test

46
Hypothesis Testing
Power of a test

47
Hypothesis Testing
Power of a test

48
Hypothesis Testing
Tests
▪ Hypothesis testing helps decide whether a population assumption (null
hypothesis, H₀) is likely true or false using sample data

▪ Tests of hypotheses (tests of significance) are broadly classified into:


Parametric tests (standard tests)
Non-parametric tests (distribution-free tests)

▪ Parametric tests assume properties about the population, such as:


➢ Normal distribution

➢ Large sample size (in many cases)

➢ Assumptions about parameters like mean and variance

49
Hypothesis Testing
Tests
▪ Non-parametric tests are used when such assumptions cannot be made

▪ Non-parametric tests:
➢ Do not depend on population parameter assumptions

➢ Often use nominal or ordinal data

➢ Usually require more observations than parametric tests for similar accuracy

▪ Major parametric tests include:


z-test | t-test | χ² (chi-square) test | F-test

50
Hypothesis Testing
Tests
z-Test:

▪ Based on the normal distribution

▪ Used for judging significance of several statistical measures, esp. mean

▪ Relevant test statistic is compared with its probable value at α

▪ Used even when binomial distribution or t-distribution is applicable

▪ Best suited for large samples when population variance is known


➢ presumption that such a distribution tends to approximate normal
distribution as ‘n’ becomes larger

51
Hypothesis Testing
Tests
z-Test:

▪ Used for:
➢ Testing a sample mean against a hypothesized mean (large samples)

➢ Comparing means of two independent samples

➢ Testing sample proportions

➢ Comparing proportions between two samples

➢ Testing significance of correlation, median, mode, etc

52
Hypothesis Testing
Tests
t-Test:

▪ Based on the t-distribution

▪ Best suited for small samples when population variance is unknown

▪ Used for:
➢ Testing significance of a sample mean

➢ Comparing means of two samples

➢ Paired t-test for related samples

➢ Testing significance of correlation coefficients

➢ Test statistic t is compared with critical t-values using degrees of freedom


53
Hypothesis Testing
Tests
Chi-Square (χ²) Test:

▪ Based on chi-square-distribution

▪ Used for:
➢ Comparing sample variance with population variance

➢ Also used (as non-parametric) for:


Goodness of fit | Test of independence

54
Hypothesis Testing
Tests
F-Test:

▪ Based on F-distribution

▪ Used for:
➢ Comparing variances of two independent samples

➢ Analysis of Variance (ANOVA) for comparing more than two means

➢ Testing multiple correlation coefficients

➢ Uses the F statistic and F-tables for hypothesis decisions

55
Hypothesis Testing
Tests

56
[Link]
Hypothesis Testing
Tests
1. Population normal, population infinite, sample size may be large or
small but variance of the population is known, Ha may be one-sided or
two-sided:

2. Population normal, population finite, sample size may be large or small


but variance of the population is known, Ha may be one-sided or two-
sided:

57
Hypothesis Testing
Tests
3. Population normal, population infinite, sample size small and variance of
the population unknown, Ha may be one-sided or two-sided:

4. Population normal, population finite, sample size small and variance of


the population unknown, and Ha may be one-sided or two-sided:

58
Hypothesis Testing
Tests
5. Population may not be normal but sample size is large, variance of the
population may be known or unknown, and Ha may be one-sided or two-
sided:

59
Hypothesis Testing

60
Hypothesis Testing

61
Hypothesis Testing

62
Hypothesis Testing

63
Hypothesis Testing
Tests
Example 1:

64
Hypothesis Testing
Tests
Example 1:

65
Hypothesis Testing
Tests
Example 1:

66
Hypothesis Testing
Tests
Example 2:

67
Hypothesis Testing
Tests
Example 2:

68
Hypothesis Testing
Tests
Example 2:

69
Task

A manufacturing process has a known mean output of 75 units with


population standard deviation 4 units. The production manager would
welcome any increase in mean output, but wants to guard against a decrease
in output. A random sample of 16 items gives a sample mean of 73.2 units.
Using a 5% level of significance, test whether the production process mean
has decreased.
Task

A specimen of steel wires drawn from a large lot has the following breaking
strengths (kg weight): 612, 605, 618, 610, 607, 614, 609, 611, 620, 604
Test, using Student’s t-test, whether the mean breaking strength of the lot
may be taken as: μ0=615 kg. Use a 5% significance level.
Task

A study of wheat yield was conducted using two independent samples of


fields from two agricultural regions. In the first sample of 120 fields, the
mean produce was found to be 305 lbs. per acre with a standard deviation of
14 lbs. In the second sample of 180 fields, the mean produce was 312 lbs.
per acre with a standard deviation of 16 lbs. Assuming that the two
populations have a common standard deviation of 15 lbs., test at the 5 per
cent level of significance whether the two samples may be considered to
have been drawn from the same population.
Hypothesis Testing
Difference Between Means
▪ Knowing whether the parameters of two populations are alike or different

▪ Example – Testing whether female workers earn less than male workers
for the same job

▪ For testing of difference between means:

Ho: μ1 = μ2

Where μ1 is the population mean of one population and μ2 is the


population mean of the second population, assuming both to be normal

73
Hypothesis Testing
Difference Between Means
1. Population variances are known or the samples happen to be large
samples:

74
Hypothesis Testing
Difference Between Means
2. Samples happen to be large but presumed to have been drawn from the
same population whose variance is known:

75
Hypothesis Testing
Difference Between Means
3. Samples happen to be small samples and population variances not
known but assumed to be equal:

76
Hypothesis Testing
Difference Between Means
Example 3:

77
Hypothesis Testing
Difference Between Means
Example 3:

78
Hypothesis Testing
Comparing Related Means
▪ Paired t-test is a way to test for comparing two related samples
➢ involving small values of n

➢ that does not require the variances of the two populations to be equal

➢ but the assumption that the two populations are normal must continue to
apply

▪ Observations in the two samples be collected in the form of what is called


matched pairs

▪ Such a test is generally considered appropriate in a before-and-after-


treatment study
79
Hypothesis Testing
Comparing Related Means

80
Hypothesis Testing
Difference Between Means
Example 4:

81
Hypothesis Testing
Difference Between Means
Example 4:

82
Hypothesis Testing
Difference Between Means
Example 4:

83
Hypothesis Testing
Difference Between Means
Example 4:

84
Hypothesis Testing
Proportions
▪ Used for qualitative data involving the presence or absence of an
attribute, such as success/failure, defective/non-defective, yes/no
outcomes

▪ Such data generally follows a binomial probability distribution

▪ In a binomial distribution:

Mean = np | Standard deviation = (npq)^0.5

p = probability of success

q = probability of failure
p + q = 1 | n = sample size
85
Hypothesis Testing
Proportions
▪ Instead of counting the number of successes, we often use the sample
proportion of successes

86
Hypothesis Testing
Proportions
▪ Formulate the Null Hypothesis (H0) and Alternative Hypothesis (Ha)

▪ Assume normal approximation to the binomial distribution

▪ Choose a level of significance (α)

▪ Construct the rejection (critical) region

▪ Compute the Z-statistic from sample data

▪ Compare calculated Z with critical value

▪ Decision: Reject H0 if Z falls in rejection region |Otherwise, do not reject H0

▪ The significance of the observed sample result is then judged using this
procedure
87
Hypothesis Testing
Proportions
Example 5:

88
Hypothesis Testing
Proportions
Example 5:

89
Hypothesis Testing
Difference between Proportions
▪ Drawn from different populations, one may be interested in knowing
whether the difference between the proportion of successes is significant
or not

▪ Start with the hypothesis that the difference between the proportion of
success in sample one and proportion of success in sample two
is due to fluctuations of random sampling

▪ Null Hypothesis:

90
Hypothesis Testing
Difference between Proportions

91
Hypothesis Testing
Difference between Proportions
Example 6:

92
Hypothesis Testing
Difference between Proportions
Example 6:

93
Hypothesis Testing
Comparing a Variance to a Hypothesized Population Variance
▪ Used when testing whether a sample variance differs significantly from a
theoretical or hypothesized population variance

▪ This test uses the Chi-Square (χ2) Test, rather than the Z-test or t-test

▪ Null Hypothesis is generally:

▪ Test statistic:

94
Hypothesis Testing
Comparing a Variance to a Hypothesized Population Variance
▪ Testing Procedure
➢ Calculate the Chi-square statistic

➢ Select the level of significance (α)

➢ Find the critical (table) value of χ2 for (n−1) degrees of freedom

➢ Compare calculated and tabulated values:


If calculated χ2 ≤ table value → Accept H0
If calculated χ2> table value → Reject H0

▪ Features of Chi-Square Distribution


Not symmetrical | Positively skewed | Values are always positive

▪ Degrees of freedom are essential for using this distribution


95
Hypothesis Testing
Comparing a Variance to a Hypothesized Population Variance
▪ Chi-square test (χ² / chi-square / Ki-square) is a widely used non-
parametric statistical test for analysing categorical data and variances

▪ It is primarily used to determine whether:


➢ observed data fits an expected distribution,

➢ two categorical variables are associated,

➢ or different populations show similar characteristics

▪ The chi-square test is commonly applied for:


➢ Goodness of fit testing

➢ Testing association/independence between attributes

➢ Testing homogeneity of populations or population variance 96


Hypothesis Testing
Comparing a Variance to a Hypothesized Population Variance
▪ As a non-parametric test, it does not require strict assumptions about
population distribution for categorical data analysis

▪ In variance analysis, the chi-square test helps determine whether a


sample originates from a population having a specified variance

▪ The test is based on the χ²-distribution, which arises from the summation
of squared quantities

▪ Since sample variance calculations involve squared deviations, they


naturally follow the chi-square distribution framework

97
Hypothesis Testing
Comparing a Variance to a Hypothesized Population Variance
▪ The distribution is not symmetrical and all the values are positive

▪ For making use of this distribution, one is required to know the degrees of
freedom since for different degrees of freedom we have different curves

▪ The smaller the number of degrees of freedom, the more skewed is the
distribution

98
Hypothesis Testing
Comparing a Variance to a Hypothesized Population Variance

Solution:

99
Hypothesis Testing
Comparing a Variance to a Hypothesized Population Variance

100
Hypothesis Testing
Comparing a Variance to a Hypothesized Population Variance

101
Hypothesis Testing
Chi-Square as a Non-Parametric Test
▪ Chi-square (χ²) is an important non-parametric statistical test requiring
minimal assumptions about the population

▪ The test mainly depends on: sample size, and degrees of freedom

▪ Chi-square is commonly used for: Goodness of fit testing | Testing


independence/association between attributes

▪ As a goodness of fit test, χ² determines how well observed data matches


a theoretical distribution such as:
Binomial distribution, Poisson distribution, Normal distribution

▪ A good fit indicates that differences between observed and expected


frequencies are mainly due to sampling fluctuations 102
Hypothesis Testing
Chi-Square as a Non-Parametric Test
▪ As a test of independence, χ² determines whether two attributes are
associated or independent

▪ Example: Testing whether a new medicine effectively controls fever

103
Hypothesis Testing
Chi-Square as a Non-Parametric Test
▪ As a test of independence, χ² determines whether two attributes are
associated or independent

▪ Degrees of freedom (d.f.) are essential in chi-square analysis

▪ For grouped frequency distributions: d.f. = n−1

▪ For contingency tables: d.f.= (c−1)(r−1)

c = number of columns | r = number of rows

104
Hypothesis Testing
Conditions for the Application of Chi-Square Test
▪ Observations recorded and used are collected on a random basis

▪ All the items in the sample must be independent

▪ No group should contain very few items, say less than 10. In case where
the frequencies are less than 10, regrouping is done by combining the
frequencies of adjoining groups so that the new frequencies become
greater than 10

▪ Some statisticians take the above number as 5, but 10 is regarded as


better by most

105
Hypothesis Testing
Conditions for the Application of Chi-Square Test
▪ The overall number of items must also be reasonably large

▪ It should normally be at least 50, howsoever small the number of groups


may be

▪ The constraints must be linear

▪ Constraints which involve linear equations in the cell frequencies of a


contingency table (i.e., equations containing no squares or higher powers
of the frequencies) are known are know as linear constraints

106
Hypothesis Testing
Steps Involved

107
Hypothesis Testing
Steps Involved

108
Hypothesis Testing
Steps Involved

109
Hypothesis Testing
Steps Involved

110
Hypothesis Testing
Yates Correction
▪ Yates’ correction was proposed by F. Yates for improving the accuracy of χ²
tests in small samples
▪ It is specifically applied to: 2 × 2 contingency tables, especially when
expected cell frequencies are small

▪ The correction is recommended when: cell frequencies are close to 5,and the
calculated χ² value is near the significance level
▪ Yates’ correction is also called the: continuity correction

▪ The correction reduces the difference between: observed frequencies, and


expected frequencies
▪ Applying the correction generally: lowers the χ² value, making the test more
conservative 111
Hypothesis Testing
Yates Correction

112
Hypothesis Testing
Questions

113
Hypothesis Testing
Questions

114
Hypothesis Testing
Questions

115
Hypothesis Testing
Questions
Two social researchers classified households into economic categories based
on independent sampling studies. Their observations are shown below:

Investigators Low Income Middle Income High Income Total


Researcher X 180 50 20 250
Researcher Y 120 130 50 300
Total 300 180 70 550

Using the chi-square test, determine whether the sampling technique of at


least one researcher appears to be defective.

116
Hypothesis Testing
Questions
A university conducted a study to determine whether participation in a
structured stress-management workshop helped students avoid severe
examination anxiety during final exams. The following data were collected:

Experienced Did Not Experience


Students Total
Severe Anxiety Severe Anxiety
Attended
42 558 600
Workshop
Did Not Attend
168 1232 1400
Workshop

Total 210 1790 2000

Using the chi-square test, examine whether the workshop was effective in
reducing severe examination anxiety at the 5% level of significance.
117
Hypothesis Testing
Comparing Variances of Two Normal Populations
▪ Used to test whether two populations have equal variances

▪ Based on the F-distribution

118
Hypothesis Testing
Comparing Variances of Two Normal Populations
▪ The larger variance is always placed in the numerator, so: F≥1

▪ Based on the F-distribution

▪ Decision Rule:
If calculated F > F-table→ Diff. in variances is significant→ Reject H0
If calculated F < F-table→ Diff. not significant→ Accept H0

▪ Assumptions of the F-Test


➢ Populations are normally distributed

➢ Samples are randomly drawn

➢ Observations are independent

➢ No measurement error exists 119


Hypothesis Testing
Comparing Variances of Two Normal Populations
▪ Purpose and Uses of F-Test
➢ Originally developed to test equality of two variances

➢ Widely used in Analysis of Variance (ANOVA)

➢ Useful for determining whether variability differences are statistically


significant

120
Hypothesis Testing
Comparing Variances of Two Normal Populations
Example 7:

121
Hypothesis Testing
Comparing Variances of Two Normal Populations
Example 7:

122
Hypothesis Testing
Comparing Variances of Two Normal Populations
Example 7:

123
Hypothesis Testing
Comparing Variances of Two Normal Populations
Example 7:

124
Hypothesis Testing
Correlation Coefficients
1. In case of simple correlation coefficient: We use t-test and calculate the
test statistic as under:

2. In case of partial correlation coefficient:

125
Hypothesis Testing
Correlation Coefficients

3. In case of multiple correlation coefficient: We use F-test and work out


the test statistic as under:

126
Hypothesis Testing
Correlation Coefficients

127
Hypothesis Testing
Limitations
▪ Hypothesis tests should not be used mechanically; they are aids to
decision-making, not decisions themselves

▪ Statistical tests support judgment, but proper interpretation of statistical


evidence is essential for intelligent conclusions

▪ A hypothesis test only indicates whether a difference is statistically


significant, not whether it is practically or scientifically important

▪ Tests do not explain the causes behind an observed difference, such as


why two sample means differ

▪ They only help determine whether a difference may be due to sampling


fluctuations or other factors, but do not identify those factors 128
Hypothesis Testing
Limitations
▪ Results of significance tests are based on probability, not absolute
certainty

▪ When a result is statistically significant, it means the difference is


probably not due to chance, not that it is proven with complete certainty

▪ Statistical inferences drawn from tests are not infallible evidence


regarding the truth of a hypothesis

▪ This limitation becomes more serious with small sample sizes, where the
likelihood of erroneous inferences is generally higher

▪ Larger samples are often needed to improve the reliability and validity of
test results 129
Hypothesis Testing
Limitations
▪ Statistical significance alone should not be the sole basis for conclusions;
subject-matter knowledge is equally important.

▪ Effective use of hypothesis testing requires combining:


Statistical techniques
Knowledge of the problem domain
Sound judgment and interpretation

▪ Overall, tests of hypotheses are valuable tools, but they have limitations
and should be applied thoughtfully, not blindly

130
Task

In a demographic survey conducted in a district hospital, records of 3,232


births showed that 1,705 were male babies and the remaining 1,527 were
female babies. It is generally hypothesized that male and female births occur
in equal proportion, that is, the probability of a male birth is 0.5 and the
probability of a female birth is [Link] a 5% level of significance, test
whether the observed data support the hypothesis that the sex ratio is
50:50, or whether the deviation from equality is statistically significant.
Task

A manufacturing company has historically found that 10% of its products are
defective. A supplier introduces a new raw material and claims that using this
material will reduce the defect rate below 10%.To verify this claim, the
company produces a random sample of 400 items using the new raw
material. Out of these, 34 items are found defective. Using a 1% level of
significance, test whether there is sufficient statistical evidence to support
the supplier’s claim that the new material reduces the proportion of
defectives.
Task

A public health department wants to study whether a heavy increase in


tobacco tax has reduced smoking in a large city. Before the tax increase, a
random sample of 500 men showed that 400 were smokers. After the tax
increase, another independent random sample of 600 men found 400
smokers. Using a 5% level of significance, test whether the observed
reduction in the proportion of smokers after the tax increase is statistically
significant.

You might also like