Re 14
Re 14
Chapter 14
Hypothesis Testing
(SPECIFIC to GENERAL)
LOGICALLY TRUE BUT REALISTICALLY NOT TRUE
-Deduction Reasoning
-Induction Reasoning
A hypothesis is an unsubstantiated assumption about the relationship between concepts and constructs; it
is conjecture. It guides the research and is formulated for testing.
Induction moves from specific facts to general, but tentative, conclusions. With the aid of probability
estimates, we can qualify our results and state the degree of confidence we have in them. Statistical
inference is an application of inductive reasoning. It allows us to reason from evidence found in the sample to
conclusions we wish to make about the population.
(SPECIFIC to GENERAL) LOGICALLY TRUE BUT REALISTICALLY NOT TRUE
MANGO BASKET
Deduction is a form of reasoning in which the conclusion must necessarily follow from the premises given.
Any hypothesis based on a deduction is only as good as the premises on which it is based. Recall that
induction and deduction were discussed in the Chapter on Research Foundations and Fundamentals. The
two hidden slides that follow are the slides that appeared in that chapter’s slide deck. Reveal them if you
want to review these concepts in more detail.
• Inductions are an inferential leap beyond the evidence presented (lots of other factors contribute to a
person’s weight, not just the consumption of dessert).
• Deductions are only as good as the premise(s) on which they are based. If one of the premises is
false (that thieves know that registered computers are hard to fence), the conclusion cannot be true.
Deductions often seem logical to us, as they reflect our beliefs and we have a hard time believing
they might not be true.
• A deduction is a hypothesis that must be tested.
Deductive reasoning, or deduction, is making an inference based on widely accepted facts or premises. .
Statistical Procedures
Inferential Statistics = estimation of population values and the testing of statistical hypotheses
Descriptive Statistics = Descriptive statistics simply describe the characteristics of the data by
giving frequencies, measures of central tendency, and dispersion
Inferential statistics includes the estimation of population values and the testing of statistical hypotheses.
Descriptive statistics simply describe the characteristics of the data by giving frequencies, measures of
central tendency, and dispersion. These concepts were discussed in Chapter 13: Collect, Prepare and
Examine Data and the Appendix Describing Data Statistically.
Under the heading inferential statistics, two topics are discussed. The first, estimation of population values,
was used with sampling (chapter 5) and it will be discussed again here.
The second, testing statistical hypotheses, is the primary subject of this chapter.
Descriptive statistics describes data (for example, a chart or graph) and inferential statistics allows you to
make predictions (“inferences”) from that data. ... This is where you can use sample data to answer research
questions
After you have detailed your hypotheses in your preliminary analysis plan, the purpose of hypothesis testing is
to determine the accuracy of your hypotheses due to the fact that you have collected a sample of data, not a
census. Exhibit 14-2 reminds you of the relationships among your design strategy, data collection activities,
preliminary analysis, and hypothesis testing.
Bayesian statistics
-Extension of classical approach
-Analysis based on sample data
-Also considers established subjective probability estimates
Although there are two approaches to hypothesis testing, the more established is the classical or sampling-
theory approach.
Classical statistics are found in all of the major statistics books and are widely used in research
applications. This approach represents an objective view of probability in which the decision making rests
totally on an analysis of available sampling data. A hypothesis is established; it is then rejected or fails to be
rejected, based on the sample data collected.
The second approach is known as Bayesian statistics, which are an extension of the classical approach. It
also uses sampling data, but it goes beyond to consider all other available information. These subjective
estimates are based on general experience rather than on specific collected data. Various decision rules are
established, cost and other estimates can be introduced, and the expected outcomes of combinations of
these elements are used to judge decision alternatives.
Bayesian statistics is a theory in the field of statistics based on the Bayesian interpretation of probability
where probability expresses a degree of belief in an event. ...
For example, if we want to find the probability of selling ice cream on a hot and sunny day, Bayes’ theorem
gives us the tools to use prior knowledge about the likelihood of selling ice cream on any other type of day
(rainy, windy, snowy etc.).
In this case, the difference is based on a census of the production vehicles and there is no sampling involved.
Since it would really be too expensive to analyze all of the manufacturer’s vehicles, we could resort to
sampling.
Assume a sample of 25 cars is randomly selected and the average miles per gallon city is calculated to be 54.
Is 51 significantly different from 54 or is it only sampling error? Hypothesis testing will answer this question.
A difference has statistical significance if there is good reason to believe the difference does not represent
random sampling fluctuations only.
Statistical significance is a determination about the null hypothesis, which suggests that the results are due
to chance alone
A difference has practical significance if it is large enough to be meaningful—to have real value to the
decision-maker.
The null hypothesis is used for testing. It is a statement that no difference exists between the parameter and
the statistic being compared to it. The parameter is a measure taken by a census of the population or a prior
measurement of a sample of the population.
Analysts usually test to determine whether there has been no change in the population of interest or whether
a real difference exists.
In the hybrid-vehicle example, the null hypothesis states that the population parameter of 50 mpg has not
changed. An alternative hypothesis holds that there has been no change in average mpg. The alternative is
the logical opposite of the null hypothesis. This is a two-tailed test. A two-tailed test is a nondirectional test to
reject the hypothesis that the sample statistic is either greater than or less than the population parameter.
A one-tailed test is a directional test of a null hypothesis that assumes the sample parameter is not the same
as the population statistic, but that the difference can be in only one direction. The other hypotheses shown
are directional.
A hypothesis is a speculation or theory, based on insufficient evidence, that lends itself to further testing and
experimentation. With further testing, a hypothesis can usually be proven to be true or false.
In the null hypothesis, there is no relationship between the two variables. In the alternative hypothesis, there
is some relationship between the two variables i.e. They are dependent upon each other. 2. Generally,
researchers and scientists try to reject or disprove the null hypothesis
The actual test begins by considering two hypotheses. They are called the null hypothesis and the alternative
hypothesis.
A null hypothesis is a type of hypothesis used in statistics that proposes that there is no difference between
certain characteristics of a population (or data-generating process).
The null hypothesis is a typical statistical theory which suggests that no statistical relationship and
significance exists in a set of given single observed variable
Salary and job satisfaction
H0: (salary ≠ JS
HA: (salary) = Job sat
The alternate hypothesis is just an alternative to the null
A null hypothesis is a statement, in which there is no relationship between two variables. An alternative
hypothesis is statement in which there is some statistical significance between two measured phenomenon
There are two options for a decision. They are “reject H 0” if the sample information favors the alternative
hypothesis or “do not reject H 0” or “decline to reject H 0” if the sample information is insufficient to reject the
null hypothesis.
Decision Rule
Take no corrective action if the analysis shows that one cannot reject the null hypothesis.
Note the language “cannot reject” rather than “accept” the null hypothesis. It is argued that a null hypothesis
can never be proved and therefore cannot be accepted.
Statistical Decisions
Exhibit 14-4
In our system of justice, the innocence of an accused person is assumed until proof of guilt beyond a
reasonable doubt can be established. In hypothesis testing, this is the null hypothesis; there should be
difference between the assumption of innocence and the outcome unless contrary evidence is furnished.
Once evidence establishes beyond a reasonable doubt that innocence can no longer be maintained, a just
conviction is required. This is equivalent to rejecting the null hypothesis and accepting the alternative
hypothesis. Incorrect decisions or errors are the other two possible outcomes. We can justly convict an
innocent person or we can acquit a guilty person.
Exhibit 14-4 compares the statistical situation to the legal system. One of two conditions exists – either the
null hypothesis is true or the alternate is true.
When a Type I error (α) is committed,
-a true null is rejected; the innocent is unjustly convicted.
-The alpha value is called the level of significance and is the probability of rejecting the true null.
A Type I error is often represented by the Greek letter alpha (α) and a Type II error by the Greek letter beta (β
). In choosing a level of probability for a test, you are actually deciding how much you want to risk committing
a Type I error—rejecting the null hypothesis when it is, in fact, true
Probability of Making a Type I Error
Since the distribution of sample means is normal, the critical values can be computed in terms of the
standardized random variable. In this example, the critical values that provide a Type I error of .05 are 46.08
and 53.92
Probability of Making A Type I Error
The manufacturer would commit a Type II error by accepting the null hypothesis when in truth the mpg had
changed. This kind of error is difficult to detect.
To illustrate, assume has actually moved to 54 from 50. Use the formula from page 499.
Using Exhibit C-1 in Appendix C, we interpolate between .35 and .36 Z scores to find the .355 Z score. The area
between the mean and Z is .1387. is the tail area, or the area below the Z and is calculated as: Use the
formula from page 499.
This shown in the Exhibit. It is the percent of the area where we would not reject the null, when in fact it was
false because the true mean was 54. There is a 36% probability of a Type II error if the is 54. The power of the
test is 1 minus the probability of committing a Type II error. In this example, the power of the test equals 64%.
In other words, we will correctly reject the false null hypothesis with a 64% probability. A power of 64% is less
than the 80% recommended by statisticians.
1. State the null hypothesis. While the researcher is usually interesting in testing a hypothesis of change
or differences, the null hypothesis is always used for statistical testing purposes.
2. Choose the statistical test. To test a hypothesis, one must choose an appropriate statistical test.
There are many tests from which to choose. Test selection is covered more later in this chapter.
3. Select the desired level of significance. The choice of the level of significance should be made before
data collection. The most common level is .05. Other levels used include .01, .10, .025, and .001. The
exact level is largely determined by how much risk one is willing to accept and the effect this choice
has on Type II risk. The larger the Type I risk, the lower the Type II risk.
4. Compute the calculated difference value. After data collection, use the formula for the appropriate
statistical test to obtain the calculated value. This can be done by hand or with a software program.
5. Obtain the critical test value. Look up the critical value in the appropriate table for that distribution.
6. Interpret the test. For most tests, if the calculated value is larger than the critical value, reject the null
hypothesis. If the critical value is larger, fail to reject the null.
Tests of Significance
Nonparametric - Parametric
Parametric tests are significance tests for data from interval or ratio scales. They are more powerful than
nonparametric tests.
Nonparametric tests are used to test hypotheses with nominal and ordinal data.
Parametric tests should be used if their assumptions are met.
Parametric statistics are based on assumptions about the distribution of population from which the sample
was taken. Nonparametric statistics are not based on assumptions, that is, the data can be collected from a
sample that does not follow a specific distribution.
Parametric tests assume a normal distribution of values, or a “bell-shaped curve.” For example, height is
roughly a normal distribution in that if you were to graph height from a group of people, one would see a
typical bell-shaped curve. ... Nonparametric tests are used in cases where parametric tests are not
appropriate.
Probability Plot
Exhibit 14-7 (top left) Exhibit 14-7 (top right) Exhibit 14-7 (bottom)
The normality of the An alternative way to look at In these panels, there is neither a straight line in
distribution may be checked in this is to plot the deviations the normal probability plot nor a random
several ways. One such tool is from the straight line. Here we distribution of points about 0 in the detrended
the normal probability plot. would expect the points to plot. This tells us that the variable is not normally
This plot compares the cluster without pattern around distributed.
observed values with those a straight line passing
expected from a normal horizontally through 0.
distribution. If the data display
the characteristics of
normality, the points will fall
within a narrow band along a
straight line. An example is
shown in the slide.
The normal probability plot
(Chambers et al., 1983) is
a graphical technique for
assessing whether or not a
data set is approximately
normally distributed. The data
are plotted against a
theoretical normal distribution
in such a way that the points
should form an approximate
straight line.
Advantages of Nonparametric Tests
Easy to understand and use
Usable with nominal data
Appropriate for ordinal data
Appropriate for non-normal population distributions
This slide lists the advantages of nonparametric tests.
Pizza preference is an example of nominal data…preference for toppings, thin vs. thick, crust, etc.
The major advantages of nonparametric statistics compared to parametric statistics are that: (1) they can be
applied to a large number of situations; (2) they can be more easily understood intuitively; (3) they can be
used with smaller sample sizes; (4) they can be used with more types of data; (5) they need fewer or less
stringent assumptions about the nature of the population distribution; (6) they are generally more robust and
not often seriously affected by extreme values in data such as outliers; (7) they have, in many cases, a high
level of asymptotic relative efficiency compared to the classical parametric tests; (8) the introduction of
jackknife, bootstrap, and other resampling techniques has increased their range of applicability; and (9) they
provide a number of supplemental or alternative tests and techniques to currently existing parametric tests.
Reference Exhibit 14-8, on the next slide, to see the recommended tests.
Parametric Tests
t-test
Z-test
The Z test or t-test is used to determine the statistical significance between a sample distribution mean
and a parameter.
A z-test, like a t-test, is a form of hypothesis testing. Where a t-test looks at two sets of data that are
different from each other — with no standard deviation or variance — a z-test views the averages of data
sets that are different from each other but have the standard deviation or variance [Link]. 22, 1443
AH
The Z distribution and t distribution differ. The t has more tail area than that found in the normal distribution.
This is a compensation for the lack of information about the population standard deviation. Although the
sample standard deviation is used as a proxy figure, the imprecision makes it necessary to go farther away
from 0 to include the percentage of values in the t distribution necessarily found in the standard normal.
When sample sizes approach 120, the sample standard deviation becomes a very good estimation of the
population standard deviation; beyond 120, the t and Z distributions are virtually identical.
A t-test is a type of inferential statistic used to determine if there is a significant difference between the
means of two groups, which may be related in certain features. The t-test is one of many tests used for the
purpose of hypothesis testing in statistics.
The one-sample t-test is used when we want to know whether our sample comes from a particular
population but we do not have full population information available to us. For instance, we may
want to know if a particular sample of college students is similar to or different from college
students in general.
One Sample Chi-Square Test Example
If the measurement scale is nominal, it is possible to use either the binomial test or the chi-square
test
In a one-sample situation, a variety of nonparametric tests may be used, depending on the measurement
scale and other conditions. If the measurement scale is nominal, it is possible to use either the binomial test
or the chi-square test. The binomial test is appropriate when the population is viewed as only two classes
such as male and female. It is also useful when the sample size is so small that the chi-square test cannot be
used.
The table illustrates the results of a survey of student interest in Metro University Dining Club. 200 students
were interviewed about their interest in joining the club. The results are classified by living arrangement. Is
there a significant difference among these students?
One-Sample Chi-Square Example
The null hypothesis states that the proportion in the population who intend to join the club is independent of
living arrangement. The alternate hypothesis states that the proportion in the population who intend to join
the club is dependent on living arrangement.
The chi-square test is used because the responses are classified into nominal categories.
Calculate the expected distribution by determining what proportion of the 200 students interviewed were in
each group. Then apply these proportions to the number who intend to join the club. Then calculate the
following:
Enter the table of critical values of X2 (Exhibit C-3) with 3 d.f., and secure a value of 7.82 at an alpha of .05.
The calculated value is greater than the critical value so the null is rejected and we conclude that intending to
join is dependent on living arrangement.
Two-Sample Parametric Tests
The Z and t-tests are frequently used parametric tests for independent samples
The Z and t-tests are frequently used parametric tests for independent samples, although the F test can also
be used.
The Z test is used with large sample sizes (exceeding 30 for both independent samples) or with smaller
samples when the data are normally distributed and population variances are known. The formula is shown in
the slide.
With small sample sizes, normally distributed populations, and the assumption of equal population
variances, the t-test is appropriate. The formula is shown in the slide.
An example is covered on the next slide.
Consider a problem facing a manager at KDL, a media firm that is evaluating account executive trainees. The
manager wishes to test the effectiveness of two methods for training new account executives. The company
selects 22 trainees who are randomly divided into two experimental groups. One receives type A and the
other type B training. The trainees are then assigned and managed without regard to the training they have
received. At the year’s end, the manager reviews the performances of these groups and finds the results
presented in the table shown in the slide.
To test whether one training method is better than the other, we will follow the standard testing procedure
shown in the next slide.
Two-Sample t-Test Example
The null hypothesis states that there is no difference is sales for group A compared group B. The alternate
hypothesis states that group A produced more sales than group B.
The t-test is chosen because the data are at least interval and the samples are independent.
The calculated value is computed as follows:
Enter Appendix Exhibit C-2 with d.f. = 20, one-tailed test, alpha = .05. The critical value is 1.725.
The calculated value is greater than the critical value so the null is rejected and we conclude that training
method A is superior.
Two-Sample Nonparametric Tests: Chi-Square
The chi-square test is appropriate for situations in which a test for differences between samples is required.
It is especially valuable for nominal data but can be used with ordinal measurements. Preparing to solve this
problem with the chi-square formula is similar to that presented earlier.
In the example in the slide, MindWriter is considering implementing a smoke-free workplace policy. It has
reason to believe that smoking may affect worker accidents. Since the company has complete records on on-
the-job accidents, a sample of workers is drawn from those who were involved in accidents during the last
year. A similar sample is drawn from among workers who had no reported accidents in the last year.
Members of both groups are interviewed to determine if each smokes on the job and whether each smoker
classifies himself or herself as a heavy or moderate smoker.
The null hypothesis states that there is no difference in distribution channel for age categories of purchasers.
The alternate hypothesis states that there is a difference in distribution channel for age categories of
purchasers.
The chi-square is chosen because the data are ordinal.
The calculated value is computed as follows:
Use the formula from page 512
The expected distribution is provided by the marginal totals of the table. The numbers of expected
observations in each cell are calculated by multiplying the two marginal totals common to a particular cell
and dividing this product by n. For example, in cell 1,1, 34 * 16/ 66 = 8.24
Enter Appendix Exhibit C-3 with d.f. = 2, and find the critical value of 5.99.
The calculated value is greater than the critical value so the null is rejected.
Exhibit 14-9
In another type of chi-square, the 2 x 2 table, a correction factor known as Yates’ correction for continuity is
applied when sample sizes are greater than 40 or when the sample is between 20 and 40 and the values of Ei
are 5 or more.
When the continuity correction is applied to the data shown in Exhibit 14-9, a chi-square of 5.25 is obtained.
The observed level of significance for this value is .02192. If the level of significance were set at .01, we would
accept the null hypothesis. However, had we calculated chi-square without the correction, the value would
have been 6.25 with an observed level of significance of .01242. The literature is in conflict over the merits of
Yates’ correction.
The Mantel-Haenszel test and the likelihood ratio also appear in Exhibit 14-9. The former is used with ordinal
data; the latter, based on maximum likelihood theory, produces results similar to Pearson’s chi-square.
Two-Related-Samples Tests
The two-related samples tests concern those situations in which persons, objects, or events are
closely matched or the phenomena are measured twice. For instance, one might compare the
consumption of husbands and wives.
Both parametric and nonparametric tests are applicable under these conditions
-Nonparametric
-Parametric
The two-related samples tests concern those situations in which persons, objects, or events are closely
matched or the phenomena are measured twice. For instance, one might compare the consumption of
husbands and wives.
Both parametric and nonparametric tests are applicable under these conditions.
Parametric
The t-test for independent samples is inappropriate here because of its assumption that observations are
independent. The problem is solved by a formula where the difference is found between each matched pair of
observations, thereby reducing the two samples to the equivalent of a one-sample case. In other words, there
are now several differences, each independent of the other, for which one can compute various statistics.
Nonparametric Tests
The McNemar test may be used with either nominal or ordinal data and is especially useful with before-after
measurement of the same subjects.
Exhibit 14-10 shows two years of Forbes sales data (in millions of dollars) from 10 companies. The next slide
illustrates the hypothesis test.
Paired-Samples t-Test Example
The null hypothesis states that there is no difference in sales data between years one and two. The alternate
hypothesis states that there is a difference.
The matched or paired-sample t-test is chosen because there are repeated measures on each company, the
data are not independent, and the measurement is ratio.
Enter Appendix Exhibit C-2 with d.f. = 9, two-tailed test, alpha = .01, and find the critical value of 3.25.
The calculated value is greater than the critical value so the null is rejected. We conclude that there is a
significant difference between the two years of sales.
SPSS Output for Paired-Samples t-Test
A computer solution to the problem is illustrated in Exhibit 14-11. Notice that an observed significance level
is printed for the calculated t value (highlighted). The observed significance level is the probability value
compared to the significance level chosen for testing and on this basis, the null hypothesis is either rejected
or not rejected.
The McNemar test may be used with either nominal or ordinal data and is especially useful with before-after
measurement of the same subjects. One can test the significance of any observed change by setting up a
fourfold table of frequencies to represent the first and second set of responses. An example is provided in the
slide.
Since A + D represents the total number of people who changed (B and C are no change responses), the null
hypothesis is that ½ (A+D) cases change in one direction and the same proportion in the other direction. The
McNemar test uses a transformation of the chi-square test.
Use formula from page 515
The minus 1 in the equation is a correction for continuity since the chi-square is a continuous distribution and
the observed frequencies represent a discrete distribution
The McNemar test may be used with either nominal or ordinal data and is especially useful with before-after
measurement of the same subjects. One can test the significance of any observed change by setting up a
fourfold table of frequencies to represent the first and second set of responses. An example is provided in the
slide.
Since A + D represents the total number of people who changed (B and C are no change responses), the null
hypothesis is that ½ (A+D) cases change in one direction and the same proportion in the other direction. The
McNemar test uses a transformation of the chi-square test.
Use formula from page 516
The minus 1 in the equation is a correction for continuity since the chi-square is a continuous distribution and
the observed frequencies represent a discrete distribution.
• The samples must be randomly selected from normal populations and the populations
should have equal variances.
• The distance from one value to its group’s mean should be independent of the distances of
other values to that mean.
Unlike the t-test, which uses sample standard deviations, ANOVA uses squared deviations of the variance to
that computation of distances of the individual data points from their own mean or from the grand mean can
be summed (recall that standard deviations sum zero).
In an ANOVA model, each group has its own mean and values that deviate from that mean. The total deviation
is the sum of the squared differences between each data point and the overall grand mean.
The total deviation of any particular data point may be partitioned into between-groups variance and within-
groups variance. The between-groups variance represents the effect of the treatment or factor. The
differences of between-groups means imply that each group was treated differently and the treatment will
appear as deviations of the sample means from the grand mean.
The within-groups variance describes the deviations of the data points within each group from the sample
mean. It is often called error.
ANOVA Example
To illustrate, consider the report about the quality of in-flight service on various carriers from the US to
Europe. Three airlines are compared. The data are shown in Exhibit 14-12. The dependent variable is service
rating and the factor is airline.
he null hypothesis states that there is no difference in the service rating score between airlines.
ANOVA and the F test is chosen because we have k independent samples, can accept the assumptions of
analysis of variance, and have interval data for the dependent variable. The significance level is .05. The
calculated F value is 28.30 (see summary table in the last slide).
Enter Appendix Exhibit C-9 with d.f. = 2, 57, and find the critical value of 3.16.
The calculated value is greater than the critical value so the null is rejected. We conclude that there is a
significant difference in flight service ratings.
Note that the p value provided in the summary table can also be used to reject the null.
Post Hoc: Scheffe’s S Multiple Comparison Procedure
Exhibit 14-14 There are several multiple comparison procedures one can use after rejecting the null with the F ratio. This
slide compares the available tests.
ANOVA Plots
Exhibit 14-15
In this exhibit, plots illustrate the comparison. The means plot shows relative differences among the three
levels of the factor. The means by standard deviations plot reveals lower variability in the opinions recorded
by the hypothetical Lufthansa and Cathay Pacific passengers. These two groups are sharply divided on the
quality of in-flight service and that is apparent in the plot.
A two-way ANOVA is used to estimate how the mean of a quantitative variable changes
according to the levels of two categorical variables. Use a two-way ANOVA when you want to
know how two independent variables, in combination, affect a dependent variable
you could use a two-way ANOVA to understand whether there is an interaction between gender
and educational level on test anxiety amongst university students, where gender
(males/females) and education level (undergraduate/postgraduate) are your independent variables,
and test anxiety is your dependent
All data are hypothetical
Exhibit 14-16
Recall that in Exhibit 14-12 data were entered for the variable seat selection: economy and business-class
travelers. If we add this factor to our model, we have a two-way analysis of variance. We can now answer
three questions:
• Do the airline and seat selections interact with respect to flight service ratings?
Exhibit 14-16, shown in the slide, tests the hypotheses for these questions. The significance level chosen is
.01. First, we consider the interaction effect of airline by seat selection. The null is accepted (p value is not
significant). But note that there are significant differences by airline and by seat selection.
When k independent samples are collected with nominal data, the chi-square is the appropriate
nonparametric technique. The Kruskal-Wallis test is appropriate for ordinal scale data or interval data tat do
not meet the F-test assumptions.
k-Related-Samples Tests
More than two levels in grouping factor
Observations are matched
Data are interval or ratio
The k-sample Anderson-Darling test is a nonparametric statistical procedure that tests the
hypothesis that the populations from which two or more groups of data were drawn are identical.
Each group should be an independent random sample from a population.
In test marketing experiments or ex post facto designs with k samples, it is often necessary to measure
subjects several times. These repeated measures are called trials.
The repeated-measures ANOVA is a special type of n-way analysis of variance. In this design, the repeated
measures of each subject are related just as they are in the related t-test where only two measures are
present. In this sense, each subject serves as its own control requiring a within-subjects variance effect to be
assessed differently than the between-groups variance in a factor like airline or seat selection.
This model is presented in Exhibit 14-18.
Summary Tables for Repeated-Measures ANOVA
Exhibit 14-18
Null hypotheses:
The F test for repeated measures is chosen because we have related trials on the dependent variable for k
samples, accept the assumptions of analysis of variance, and have interval data.
The critical test values come from Appendix Exhibit C-9; they are 3.16 and 4.01.
Based on the results, all three null hypotheses are rejected. There are significant differences in all three
cases.
Repeated Measures ANOVA Plot