0% found this document useful (0 votes)
8 views8 pages

Understanding Type I and II Errors in Testing

The document discusses errors in hypothesis testing, specifically Type I and Type II errors, and factors affecting the power of a test, including effect size, sample size, and significance level. It also covers testing hypotheses about the difference between two independent and dependent groups, detailing assumptions and potential issues with dependent-samples designs. Overall, it emphasizes the importance of understanding these concepts to accurately interpret statistical results.

Uploaded by

panwarsakshi2004
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views8 pages

Understanding Type I and II Errors in Testing

The document discusses errors in hypothesis testing, specifically Type I and Type II errors, and factors affecting the power of a test, including effect size, sample size, and significance level. It also covers testing hypotheses about the difference between two independent and dependent groups, detailing assumptions and potential issues with dependent-samples designs. Overall, it emphasizes the importance of understanding these concepts to accurately interpret statistical results.

Uploaded by

panwarsakshi2004
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Unit II

Errors in Hypothesis Testing


There are two “states of nature”: Either the null hypothesis, Ho, is true or it is
false.
Similarly, there are two possible decisions: we can reject the null hypothesis
or we can retain it.

Type-I Error

rejection of a true null hypothesis


It means concluding that results are statistically significant when, in reality,
they came about purely by chance or because of unrelated factors.
The probability of a Type I error is indicated by the Greek letter 𝜶 (alpha)

Type-II Error

retention of a false null hypothesis


It means failing to conclude there was an effect when there actually was.
The probability of committing a Type II error is indicated by the Greek letter
𝜷 (beta)
Power of test
the probability of rejecting a false null hypothesis
1−𝜷

The value of (1 − 𝛽) is called the power of the test.


Because 𝛽 and the power of a test are complementary, any condition that
decreases 𝛽 increases the power of the test, and vice versa.

Factors Affecting Power of Test


1. Difference between the True Population Mean and the Hypothesized Mean
(Size of Effect): the greater the difference between 𝜇true and 𝜇hyp, the
more powerful is the test.
2. Sample Size: Increasing the sample size increases the power of the test.
Larger samples provide more information about the population, reducing the
standard error and making it easier to detect a true effect
3. Variability of the Measure: Decreasing variability (standard deviation)
increases the power of the test (inversely proportional). Less variability
within the data means that any true effect is easier to detect because the
signal-to-noise ratio is higher
4. Level of Significance (𝜶 ): reducing the risk of Type I error increases the risk
of committing a Type II error and thus reduces the power of the test.
Increasing the significance level increases the power of the test. A higher
(alpha) means a higher threshold for rejecting the null hypothesis, which
increases the chance of detecting an effect if it exists.
5. One-Tailed versus Two-Tailed Tests: the probability of rejecting a false null
hypothesis is greater for a one-tailed test than for a two-tailed test (but only
when the direction specified by HA is correct). A one-tailed test does not
split the significance level across two tails, making it easier to reject the null
hypothesis for an effect in the specified direction.
6. Effect Size (d): Larger effect sizes increase the power of the test. When the
true difference between groups or the true effect is larger, it is easier to
detect.
Testing Hypotheses about the Difference between
Two Independent Groups
Independent Samples- samples drawn so that the selection of elements in
one sample is in no way influenced by the selection of elements in the other,
and vice versa
The Null and Alternative Hypotheses

Random Sampling Distribution of the Difference between two sample


means: a theoretical relative frequency distribution of all the values of (X bar
− Y bar) that would be obtained by chance from an infinite number of
samples of a particular size drawn from two populations, 𝜇X and 𝜇Y
Properties of the Sampling Distribution of the Difference between Means
- the mean of the random sampling distribution of the difference between
pairs of sample means, 𝜇X−Y , is the same as the difference between the
two population means, 𝜇X − 𝜇Y.
- sampling distribution of (X − Y) will be normally distributed when the two
populations are normally distributed and that the sampling distribution tends
toward normal even when the two populations are not normal

- The standard deviation of the sampling distribution of is called


the standard error of the difference between two means, and its symbol is

However, we usually do not know the population standard deviations and so


must estimate them from the samples.
Thus we calculate estimated standard error of the difference between two
means
When sample size is not equal

When sample size is equal

Directional (one-tailed) test the alternative hypothesis states that 𝜇X differs


from 𝜇Y in one particular direction.

Sample Size in Inference about Two Means


If 𝜎X and 𝜎Y are equal, when nX = nY the result is a smaller value of

.
The advantage of a smaller standard error of difference between 2 means is
that, if there is a difference between 𝜇X and 𝜇Y, the probability of rejecting
H0 is increased.
Large samples increase the probability of detecting a difference when a
difference exists
Power of a Test: The probability of correctly rejecting a false null hypothesis.
Higher power means a higher likelihood of detecting a true effect.
Level of Significance (α): The risk of rejecting a true null hypothesis (Type I
error). Common values are 0.05 (5%) and 0.01 (1%).
Effect Size (d): The magnitude of the difference between the means of two
groups. A larger effect size makes it easier to detect a difference.
Sample Size (n): The number of subjects in each group. Larger sample sizes
generally provide more power to detect an effect.

Assumptions Associated with Inference about the


Difference between Two Independent Means
Each sample is drawn at random from its respective population.
The two random samples must be independently selected.
Independence within groups: The observations within each group
should be independent of each other. This means that the value of one
observation does not affect or influence the value of another
observation within the same group.
Independence between groups: The two groups being compared must
be independent of each other.
Samples are drawn with replacement.
The distributions of the populations from which the samples are drawn
should be approximately normally distributed
Homogeneity of Variance: The variances of the populations from which the
two samples are drawn should be equal (or nearly equal). When this
assumption is violated, the results of the standard t-test can be inaccurate.

Testing for a Difference between Two Dependent


(Correlated) Groups
dependent-samples design- a study in which measurements in one sample
are related to measurements in the other sample(s)
repeated-measures design- a study in which observations are measured on
the same subject under two (or more) conditions
The potential advantage of this design is that it can reduce differences
between the two groups of scores due to random sampling.
matched-subjects design- a study in which subjects from two (or more)
samples are paired (matched) by the experimenter on an extraneous
variable
matched-pairs investigations- a study in which subjects from two groups
are naturally paired
For dependent means, the ideal model supposes that there exists a
population of paired observations (through matching or repeated measures)
and that a sample is selected at random from this population.
It also presupposes that we know 𝜌 (rho), the correlation between pairs of
observations in the population.

df = n − 1, where n is the number of pairs of scores.

An Alternative Approach to the Problem of Two Dependent


Means
This method focuses on the characteristics of the sampling distribution of
differences between the paired X and Y scores.

D symbol for a difference score, X − Y

symbol for the mean of a sample of difference scores


The hypothesis may be stated as Ho: 𝜇D = 0
this transforms the test from a two-sample test (of X and Y scores) to a one-
sample test (of difference scores)
symbol for the estimated population standard deviation of D

Power
power of a test- the probability of rejecting a false null hypothesis
An increase in sample size increases power by reducing the standard error
of the mean (s)
The dependent-samples design makes it possible to reduce the standard
error of the difference between the means by controlling the influence of
extraneous variables.

Assumptions When Testing a Hypothesis about the


Difference between Two Dependent Means
Paired Observations: Each observation in one sample has a corresponding
observation in the other sample. The data must be in pairs where each pair
consists of two related or matched observations.
the assumption of homogeneity of variance is not required for the test
between two dependent means
our sample has been drawn randomly with replacement and that the
sampling distribution of D is normally distributed
When sample size is large (>30), the central limit theorem tells us that the
sampling distribution of difference between two means should approximate
the normal curve regardless of what the original population of scores looks
like.

Problems with Using the Dependent-Samples Design


Problem 1: Order Effects
- when making repeated measurements on the same subjects, exposure to the
treatment condition assigned first, changes performance on the treatment
condition assigned second
- This tends to increase the standard error and, consequently, to decrease the
chance of detecting a difference when one exists.
Problem 2: Repeated-Measures Design with Non-random Assignment
- If a design uses repeated observations on the same subjects but the
assignment of the order of the treatment condition is not random, it may be
difficult to interpret the outcome.
- Any order effect will bias the comparison.
Problem 3: Sampling Variation in Matched-Subjects Designs
- In general, when pairing is on the basis of a variable importantly related to the
performance of the subjects, the correlation will be high and the reduction in

will consequently be great.

- A reduction in increases the power of the test.

You might also like