0% found this document useful (0 votes)
15 views22 pages

Re 14

This document discusses hypothesis testing, focusing on the distinction between inductive and deductive reasoning in formulating and testing hypotheses. It outlines the processes of statistical significance, including classical and Bayesian approaches, and the concepts of null and alternative hypotheses. Additionally, it explains the implications of Type I and Type II errors in the context of statistical decision-making.

Uploaded by

BESHOO
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
15 views22 pages

Re 14

This document discusses hypothesis testing, focusing on the distinction between inductive and deductive reasoning in formulating and testing hypotheses. It outlines the processes of statistical significance, including classical and Bayesian approaches, and the concepts of null and alternative hypotheses. Additionally, it explains the implications of Type I and Type II errors in the context of statistical decision-making.

Uploaded by

BESHOO
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Stage 4: Hypothesis Testing

Chapter 14

of an appropriate test of statistical significance.


How to interpret the various test statistics.

Hypothesis Testing
(SPECIFIC to GENERAL)
LOGICALLY TRUE BUT REALISTICALLY NOT TRUE

-Deduction Reasoning
-Induction Reasoning
A hypothesis is an unsubstantiated assumption about the relationship between concepts and constructs; it
is conjecture. It guides the research and is formulated for testing.
Induction moves from specific facts to general, but tentative, conclusions. With the aid of probability
estimates, we can qualify our results and state the degree of confidence we have in them. Statistical
inference is an application of inductive reasoning. It allows us to reason from evidence found in the sample to
conclusions we wish to make about the population.
(SPECIFIC to GENERAL) LOGICALLY TRUE BUT REALISTICALLY NOT TRUE
MANGO BASKET
Deduction is a form of reasoning in which the conclusion must necessarily follow from the premises given.
Any hypothesis based on a deduction is only as good as the premises on which it is based. Recall that
induction and deduction were discussed in the Chapter on Research Foundations and Fundamentals. The
two hidden slides that follow are the slides that appeared in that chapter’s slide deck. Reveal them if you
want to review these concepts in more detail.

Reasoning and Hypotheses

Inductions are an inferential leap from the evidence presented.


Statistical inference is the process of using data analysis to infer properties of an underlying distribution of
probability. Inferential statistical analysis infers properties of a population, for example by testing hypotheses
and deriving estimates
Induction starts with drawing a conclusion (desserts cause weight gain) from one or more particular facts or
pieces of evidence (people who eat dessert are heavier than those who don’t eat dessert).

• Inductions are an inferential leap beyond the evidence presented (lots of other factors contribute to a
person’s weight, not just the consumption of dessert).

• An Induction is a hypothesis that must be tested.

Reasoning and Hypotheses


Deductions are only as good as the premises on which they are based.
ALL MEN ARE TALL. -----THEREFORE MR. AHMAD IS TALL
Deduction starts with one or more premises (thieves steal to fence the items they steal for money; thieves
know they can’t fence registered computers;) and the conclusion must follow from the premise (Thieves
won’t steal registered computers).

• Deductions are only as good as the premise(s) on which they are based. If one of the premises is
false (that thieves know that registered computers are hard to fence), the conclusion cannot be true.
Deductions often seem logical to us, as they reflect our beliefs and we have a hard time believing
they might not be true.
• A deduction is a hypothesis that must be tested.

Deductive reasoning, or deduction, is making an inference based on widely accepted facts or premises. .

Statistical Procedures

Inferential Statistics = estimation of population values and the testing of statistical hypotheses
Descriptive Statistics = Descriptive statistics simply describe the characteristics of the data by
giving frequencies, measures of central tendency, and dispersion
Inferential statistics includes the estimation of population values and the testing of statistical hypotheses.
Descriptive statistics simply describe the characteristics of the data by giving frequencies, measures of
central tendency, and dispersion. These concepts were discussed in Chapter 13: Collect, Prepare and
Examine Data and the Appendix Describing Data Statistically.
Under the heading inferential statistics, two topics are discussed. The first, estimation of population values,
was used with sampling (chapter 5) and it will be discussed again here.
The second, testing statistical hypotheses, is the primary subject of this chapter.
Descriptive statistics describes data (for example, a chart or graph) and inferential statistics allows you to
make predictions (“inferences”) from that data. ... This is where you can use sample data to answer research
questions

Hypothesis Testing and the Research Process

After you have detailed your hypotheses in your preliminary analysis plan, the purpose of hypothesis testing is
to determine the accuracy of your hypotheses due to the fact that you have collected a sample of data, not a
census. Exhibit 14-2 reminds you of the relationships among your design strategy, data collection activities,
preliminary analysis, and hypothesis testing.

The “Ah-Ha” Moment


When researchers sift through the chaos and find what matters.
While exploratory data analysis revealed some relationships worth testing, those results do not carry the
weight of results revealed when the hypothesis is tested for significance.
Although there are two approaches to hypothesis testing, the more established is the classical or sampling-
theory approach. Classical statistics are found in all of the major statistics books and are widely used in
research applications. This approach represents an objective view of probability in which the decision making
rests totally on an analysis of available sampling data. A hypothesis is established; it is then rejected or fails
to be rejected, based on the sample data collected.
Statistical Significance
Following a classical statistics approach, we accept or reject a hypothesis on the basis of data collected from
the sample alone. Because any sample will almost surely vary somewhat from its population, we must judge
whether the differences are statistically significant or insignificant. A difference has statistical significance if
there is good reason to believe the difference does not represent random sampling fluctuations only
Statistical significance is a determination about the null hypothesis, which suggests that the results are due
to chance alone

Approaches to Hypothesis Testing


Classical statistics
-Objective view of probability
-Established hypothesis is rejected or fails to be rejected
-Analysis based on sample data

Bayesian statistics
-Extension of classical approach
-Analysis based on sample data
-Also considers established subjective probability estimates
Although there are two approaches to hypothesis testing, the more established is the classical or sampling-
theory approach.
Classical statistics are found in all of the major statistics books and are widely used in research
applications. This approach represents an objective view of probability in which the decision making rests
totally on an analysis of available sampling data. A hypothesis is established; it is then rejected or fails to be
rejected, based on the sample data collected.
The second approach is known as Bayesian statistics, which are an extension of the classical approach. It
also uses sampling data, but it goes beyond to consider all other available information. These subjective
estimates are based on general experience rather than on specific collected data. Various decision rules are
established, cost and other estimates can be introduced, and the expected outcomes of combinations of
these elements are used to judge decision alternatives.
Bayesian statistics is a theory in the field of statistics based on the Bayesian interpretation of probability
where probability expresses a degree of belief in an event. ...
For example, if we want to find the probability of selling ice cream on a hot and sunny day, Bayes’ theorem
gives us the tools to use prior knowledge about the likelihood of selling ice cream on any other type of day
(rainy, windy, snowy etc.).

Significance & Hypotheses


Statistical Significance
Practical Significance
Consider this example: The hybrid Toyota Prius, shown above, inspires a cult-like devotion from its drivers,
maintaining satisfaction rates at 98 percent. Let’s say that the Prius has maintained an average of about 51
miles per gallon city with a standard deviation of 10 miles per gallon and researchers discover by analyzing all
production vehicles that the miles per gallon is now 50.

• Is the difference statistically significant? If 51 significantly different than 50?

In this case, the difference is based on a census of the production vehicles and there is no sampling involved.
Since it would really be too expensive to analyze all of the manufacturer’s vehicles, we could resort to
sampling.
Assume a sample of 25 cars is randomly selected and the average miles per gallon city is calculated to be 54.
Is 51 significantly different from 54 or is it only sampling error? Hypothesis testing will answer this question.
A difference has statistical significance if there is good reason to believe the difference does not represent
random sampling fluctuations only.
Statistical significance is a determination about the null hypothesis, which suggests that the results are due
to chance alone
A difference has practical significance if it is large enough to be meaningful—to have real value to the
decision-maker.
The null hypothesis is used for testing. It is a statement that no difference exists between the parameter and
the statistic being compared to it. The parameter is a measure taken by a census of the population or a prior
measurement of a sample of the population.
Analysts usually test to determine whether there has been no change in the population of interest or whether
a real difference exists.
In the hybrid-vehicle example, the null hypothesis states that the population parameter of 50 mpg has not
changed. An alternative hypothesis holds that there has been no change in average mpg. The alternative is
the logical opposite of the null hypothesis. This is a two-tailed test. A two-tailed test is a nondirectional test to
reject the hypothesis that the sample statistic is either greater than or less than the population parameter.
A one-tailed test is a directional test of a null hypothesis that assumes the sample parameter is not the same
as the population statistic, but that the difference can be in only one direction. The other hypotheses shown
are directional.

Null vs. Alternative Hypotheses

A hypothesis is a speculation or theory, based on insufficient evidence, that lends itself to further testing and
experimentation. With further testing, a hypothesis can usually be proven to be true or false.
In the null hypothesis, there is no relationship between the two variables. In the alternative hypothesis, there
is some relationship between the two variables i.e. They are dependent upon each other. 2. Generally,
researchers and scientists try to reject or disprove the null hypothesis
The actual test begins by considering two hypotheses. They are called the null hypothesis and the alternative
hypothesis.
A null hypothesis is a type of hypothesis used in statistics that proposes that there is no difference between
certain characteristics of a population (or data-generating process).
The null hypothesis is a typical statistical theory which suggests that no statistical relationship and
significance exists in a set of given single observed variable
Salary and job satisfaction
H0: (salary ≠ JS
HA: (salary) = Job sat
The alternate hypothesis is just an alternative to the null
A null hypothesis is a statement, in which there is no relationship between two variables. An alternative
hypothesis is statement in which there is some statistical significance between two measured phenomenon
There are two options for a decision. They are “reject H 0” if the sample information favors the alternative
hypothesis or “do not reject H 0” or “decline to reject H 0” if the sample information is insufficient to reject the
null hypothesis.

Two-Tailed Test of Significance


Exhibit 14-3, left side This is an illustration of a two-tailed test. It is a non- directional test.
In statistics, a two-tailed test is a method in which the critical area of a distribution is two-sided and tests
whether a sample is greater or less than a range of values. ... By convention two-tailed tests are used to
determine significance at the 5% level, meaning each side of the distribution is cut at 2.5%.
What is a two-tailed test? First let’s start with the meaning of a two-tailed test. If you are using a significance
level of 0.05, a two-tailed test allots half of your alpha to testing the statistical significance in one direction
and half of your alpha to testing statistical significance in the other direction. This means that .025 is in each
tail of the distribution of your test statistic. When using a two-tailed test, regardless of the direction of the
relationship you hypothesize, you are testing for the possibility of the relationship in both directions. For
example, we may wish to compare the mean of a sample to a given value x using a t-test. Our null hypothesis
is that the mean is equal to x. A two-tailed test will test both if the mean is significantly greater than x and if
the mean significantly less than x. The mean is considered significantly different from x if the test statistic is in
the top 2.5% or bottom 2.5% of its probability distribution, resulting in a p-value less than 0.05.
The z-test is also a hypothesis test in which the z-statistic follows a normal distribution. The z-test is best
used for greater-than-30 samples because, under the central limit theorem, as the number of samples gets
larger, the samples are considered to be approximately normally distributed.
If a z-score is equal to 0, it is on the mean. A positive z-score indicates the raw score is higher than the mean
average. For example, if a z-score is equal to +1, it is 1 standard deviation above the mean. A negative z-score
reveals the raw score is below the mean average.
One-Tailed Test of Significance

This is an illustration of a one-tailed, or directional, test.


What is a one-tailed test?
Next, let’s discuss the meaning of a one-tailed test. If you are using a significance level of .05, a one-tailed
test allots all of your alpha to testing the statistical significance in the one direction of interest. This means
that .05 is in one tail of the distribution of your test statistic. When using a one-tailed test, you are testing for
the possibility of the relationship in one direction and completely disregarding the possibility of a relationship
in the other direction

Decision Rule

Take no corrective action if the analysis shows that one cannot reject the null hypothesis.
Note the language “cannot reject” rather than “accept” the null hypothesis. It is argued that a null hypothesis
can never be proved and therefore cannot be accepted.
Statistical Decisions
Exhibit 14-4
In our system of justice, the innocence of an accused person is assumed until proof of guilt beyond a
reasonable doubt can be established. In hypothesis testing, this is the null hypothesis; there should be
difference between the assumption of innocence and the outcome unless contrary evidence is furnished.
Once evidence establishes beyond a reasonable doubt that innocence can no longer be maintained, a just
conviction is required. This is equivalent to rejecting the null hypothesis and accepting the alternative
hypothesis. Incorrect decisions or errors are the other two possible outcomes. We can justly convict an
innocent person or we can acquit a guilty person.
Exhibit 14-4 compares the statistical situation to the legal system. One of two conditions exists – either the
null hypothesis is true or the alternate is true.
When a Type I error (α) is committed,
-a true null is rejected; the innocent is unjustly convicted.
-The alpha value is called the level of significance and is the probability of rejecting the true null.
A Type I error is often represented by the Greek letter alpha (α) and a Type II error by the Greek letter beta (β
). In choosing a level of probability for a test, you are actually deciding how much you want to risk committing
a Type I error—rejecting the null hypothesis when it is, in fact, true
Probability of Making a Type I Error

Exhibit 14-5 (top)


Assume the hybrid car manufacturer’s problem is complicated by a consumer testing agency’s assertion that
the average mpg has changed.
Assume the population mean is 50 mpg, the standard deviation is 10 mpg, and the size of the sample is 25
vehicles.
With this information, one can calculate the standard error of the mean (the standard deviation of the
distribution of sample means). This hypothetical distribution is pictured in Exhibit 17-4. The standard error of
the mean is calculated to be 2 mpg.
If the decision is to reject Ho with a 95% confidence interval (alpha = .05), a Type I error of .025 in each tail is
accepted (assuming a two-tailed test).
The regions of rejection are indicated by green shaded areas. The area between is the region of acceptance.
Critical Values

Since the distribution of sample means is normal, the critical values can be computed in terms of the
standardized random variable. In this example, the critical values that provide a Type I error of .05 are 46.08
and 53.92
Probability of Making A Type I Error

Exhibit 14-5 (bottom)


In this diagram, the manufacturer is interested only in increases in mpg and uses a one-tailed alternate
hypothesis. In this case, the entire region of rejection is in the upper tail of the distribution. One can accept a
5% alpha risk and compute a new critical value. Use the formula from page 435.

Factors Affecting Probability of Committing a β Error


True value of parameter
Alpha level selected
One or two-tailed test used
Sample standard deviation
Sample size
Type II error is difficult to detect and the probability of committing a Type II error depends on the five factors
listed in the slide. An illustration is provided on the next slide.
With a Type II error (),
-one fails to reject a false null hypothesis; the result is an unjust acquittal, with the guilty person going free.
-The beta value I the probability of failing to reject a false null hypothesis.
-A type II error is also known as a false negative and occurs when a researcher fails to reject a null hypothesis
which is really false
Probability of Making A Type II Error

The manufacturer would commit a Type II error by accepting the null hypothesis when in truth the mpg had
changed. This kind of error is difficult to detect.
To illustrate, assume  has actually moved to 54 from 50. Use the formula from page 499.
Using Exhibit C-1 in Appendix C, we interpolate between .35 and .36 Z scores to find the .355 Z score. The area
between the mean and Z is .1387.  is the tail area, or the area below the Z and is calculated as: Use the
formula from page 499.
This shown in the Exhibit. It is the percent of the area where we would not reject the null, when in fact it was
false because the true mean was 54. There is a 36% probability of a Type II error if the  is 54. The power of the
test is 1 minus the probability of committing a Type II error. In this example, the power of the test equals 64%.
In other words, we will correctly reject the false null hypothesis with a 64% probability. A power of 64% is less
than the 80% recommended by statisticians.

Statistical Testing Procedures


Stages
Testing for statistical significance follows a relatively well-defined pattern.

1. State the null hypothesis. While the researcher is usually interesting in testing a hypothesis of change
or differences, the null hypothesis is always used for statistical testing purposes.

2. Choose the statistical test. To test a hypothesis, one must choose an appropriate statistical test.
There are many tests from which to choose. Test selection is covered more later in this chapter.

3. Select the desired level of significance. The choice of the level of significance should be made before
data collection. The most common level is .05. Other levels used include .01, .10, .025, and .001. The
exact level is largely determined by how much risk one is willing to accept and the effect this choice
has on Type II risk. The larger the Type I risk, the lower the Type II risk.

4. Compute the calculated difference value. After data collection, use the formula for the appropriate
statistical test to obtain the calculated value. This can be done by hand or with a software program.

5. Obtain the critical test value. Look up the critical value in the appropriate table for that distribution.

6. Interpret the test. For most tests, if the calculated value is larger than the critical value, reject the null
hypothesis. If the critical value is larger, fail to reject the null.

Tests of Significance
Nonparametric - Parametric
Parametric tests are significance tests for data from interval or ratio scales. They are more powerful than
nonparametric tests.
Nonparametric tests are used to test hypotheses with nominal and ordinal data.
Parametric tests should be used if their assumptions are met.
Parametric statistics are based on assumptions about the distribution of population from which the sample
was taken. Nonparametric statistics are not based on assumptions, that is, the data can be collected from a
sample that does not follow a specific distribution.
Parametric tests assume a normal distribution of values, or a “bell-shaped curve.” For example, height is
roughly a normal distribution in that if you were to graph height from a group of people, one would see a
typical bell-shaped curve. ... Nonparametric tests are used in cases where parametric tests are not
appropriate.

Assumptions for Using Parametric Tests


Independent observations - Normal distribution -Equal variances -Interval or ratio scales
The assumptions for parametric tests include the following:
-The observations must be independent – that is, the selection of any one case should not affect the chances
for any other case to be included in the sample.
-The observations should be drawn from normally distributed populations.
-These populations should have equal variances.
-The measurement scales should be at least interval so that arithmetic operations can be used with them.

Probability Plot

Exhibit 14-7 (top left) Exhibit 14-7 (top right) Exhibit 14-7 (bottom)
The normality of the An alternative way to look at In these panels, there is neither a straight line in
distribution may be checked in this is to plot the deviations the normal probability plot nor a random
several ways. One such tool is from the straight line. Here we distribution of points about 0 in the detrended
the normal probability plot. would expect the points to plot. This tells us that the variable is not normally
This plot compares the cluster without pattern around distributed.
observed values with those a straight line passing
expected from a normal horizontally through 0.
distribution. If the data display
the characteristics of
normality, the points will fall
within a narrow band along a
straight line. An example is
shown in the slide.
The normal probability plot
(Chambers et al., 1983) is
a graphical technique for
assessing whether or not a
data set is approximately
normally distributed. The data
are plotted against a
theoretical normal distribution
in such a way that the points
should form an approximate
straight line.
Advantages of Nonparametric Tests
Easy to understand and use
Usable with nominal data
Appropriate for ordinal data
Appropriate for non-normal population distributions
This slide lists the advantages of nonparametric tests.
Pizza preference is an example of nominal data…preference for toppings, thin vs. thick, crust, etc.
The major advantages of nonparametric statistics compared to parametric statistics are that: (1) they can be
applied to a large number of situations; (2) they can be more easily understood intuitively; (3) they can be
used with smaller sample sizes; (4) they can be used with more types of data; (5) they need fewer or less
stringent assumptions about the nature of the population distribution; (6) they are generally more robust and
not often seriously affected by extreme values in data such as outliers; (7) they have, in many cases, a high
level of asymptotic relative efficiency compared to the classical parametric tests; (8) the introduction of
jackknife, bootstrap, and other resampling techniques has increased their range of applicability; and (9) they
provide a number of supplemental or alternative tests and techniques to currently existing parametric tests.

How to Select a Test

Reference Exhibit 14-8, on the next slide, to see the recommended tests.

Recommended Statistical Techniques

What is a K-sample test?


The two-sample t-test (also known as the independent samples t-test) is a method used to test whether the
unknown population means of two groups are equal
The k-sample Anderson-Darling test is a nonparametric statistical procedure that tests the hypothesis
that the populations from which two or more groups of data were drawn are identical. Each group should
be an independent random sample from a population. ... Unstructured data can often be simpler to analyze
Questions Answered by One-Sample Tests
Difference between observed and expected frequencies?
Difference between observed and expected proportions?
Significant difference between some measure of central tendency and the population parameter?
What is the one-sample t-test?
The one-sample t-test is a statistical hypothesis test used to determine whether an unknown population
mean is different from a specific value.
When can I use the test?
You can use the test for continuous data. Your data should be a random sample from a normal population.
What if my data isn’t nearly normally distributed?
If your sample sizes are very small, you might not be able to test for normality. You might need to rely on your
understanding of the data. When you cannot safely assume normality, you can perform a nonparametric test
that doesn’t assume normality.

Parametric Tests
t-test

Z-test
The Z test or t-test is used to determine the statistical significance between a sample distribution mean
and a parameter.
A z-test, like a t-test, is a form of hypothesis testing. Where a t-test looks at two sets of data that are
different from each other — with no standard deviation or variance — a z-test views the averages of data
sets that are different from each other but have the standard deviation or variance [Link]. 22, 1443
AH
The Z distribution and t distribution differ. The t has more tail area than that found in the normal distribution.
This is a compensation for the lack of information about the population standard deviation. Although the
sample standard deviation is used as a proxy figure, the imprecision makes it necessary to go farther away
from 0 to include the percentage of values in the t distribution necessarily found in the standard normal.
When sample sizes approach 120, the sample standard deviation becomes a very good estimation of the
population standard deviation; beyond 120, the t and Z distributions are virtually identical.
A t-test is a type of inferential statistic used to determine if there is a significant difference between the
means of two groups, which may be related in certain features. The t-test is one of many tests used for the
purpose of hypothesis testing in statistics.

One-Sample t-Test Example

The one-sample t-test is used when we want to know whether our sample comes from a particular
population but we do not have full population information available to us. For instance, we may
want to know if a particular sample of college students is similar to or different from college
students in general.
One Sample Chi-Square Test Example

If the measurement scale is nominal, it is possible to use either the binomial test or the chi-square
test
In a one-sample situation, a variety of nonparametric tests may be used, depending on the measurement
scale and other conditions. If the measurement scale is nominal, it is possible to use either the binomial test
or the chi-square test. The binomial test is appropriate when the population is viewed as only two classes
such as male and female. It is also useful when the sample size is so small that the chi-square test cannot be
used.
The table illustrates the results of a survey of student interest in Metro University Dining Club. 200 students
were interviewed about their interest in joining the club. The results are classified by living arrangement. Is
there a significant difference among these students?
One-Sample Chi-Square Example

The null hypothesis states that the proportion in the population who intend to join the club is independent of
living arrangement. The alternate hypothesis states that the proportion in the population who intend to join
the club is dependent on living arrangement.
The chi-square test is used because the responses are classified into nominal categories.
Calculate the expected distribution by determining what proportion of the 200 students interviewed were in
each group. Then apply these proportions to the number who intend to join the club. Then calculate the
following:
Enter the table of critical values of X2 (Exhibit C-3) with 3 d.f., and secure a value of 7.82 at an alpha of .05.
The calculated value is greater than the critical value so the null is rejected and we conclude that intending to
join is dependent on living arrangement.
Two-Sample Parametric Tests

The Z and t-tests are frequently used parametric tests for independent samples

The Z and t-tests are frequently used parametric tests for independent samples, although the F test can also
be used.
The Z test is used with large sample sizes (exceeding 30 for both independent samples) or with smaller
samples when the data are normally distributed and population variances are known. The formula is shown in
the slide.
With small sample sizes, normally distributed populations, and the assumption of equal population
variances, the t-test is appropriate. The formula is shown in the slide.
An example is covered on the next slide.

Two-Sample t-Test Example

Consider a problem facing a manager at KDL, a media firm that is evaluating account executive trainees. The
manager wishes to test the effectiveness of two methods for training new account executives. The company
selects 22 trainees who are randomly divided into two experimental groups. One receives type A and the
other type B training. The trainees are then assigned and managed without regard to the training they have
received. At the year’s end, the manager reviews the performances of these groups and finds the results
presented in the table shown in the slide.
To test whether one training method is better than the other, we will follow the standard testing procedure
shown in the next slide.
Two-Sample t-Test Example

The null hypothesis states that there is no difference is sales for group A compared group B. The alternate
hypothesis states that group A produced more sales than group B.
The t-test is chosen because the data are at least interval and the samples are independent.
The calculated value is computed as follows:
Enter Appendix Exhibit C-2 with d.f. = 20, one-tailed test, alpha = .05. The critical value is 1.725.
The calculated value is greater than the critical value so the null is rejected and we conclude that training
method A is superior.
Two-Sample Nonparametric Tests: Chi-Square

The chi-square test is appropriate for situations in which a test for differences between samples is required.
It is especially valuable for nominal data but can be used with ordinal measurements. Preparing to solve this
problem with the chi-square formula is similar to that presented earlier.
In the example in the slide, MindWriter is considering implementing a smoke-free workplace policy. It has
reason to believe that smoking may affect worker accidents. Since the company has complete records on on-
the-job accidents, a sample of workers is drawn from those who were involved in accidents during the last
year. A similar sample is drawn from among workers who had no reported accidents in the last year.
Members of both groups are interviewed to determine if each smokes on the job and whether each smoker
classifies himself or herself as a heavy or moderate smoker.

Two-Sample Chi-Square Example

The null hypothesis states that there is no difference in distribution channel for age categories of purchasers.
The alternate hypothesis states that there is a difference in distribution channel for age categories of
purchasers.
The chi-square is chosen because the data are ordinal.
The calculated value is computed as follows:
Use the formula from page 512
The expected distribution is provided by the marginal totals of the table. The numbers of expected
observations in each cell are calculated by multiplying the two marginal totals common to a particular cell
and dividing this product by n. For example, in cell 1,1, 34 * 16/ 66 = 8.24
Enter Appendix Exhibit C-3 with d.f. = 2, and find the critical value of 5.99.
The calculated value is greater than the critical value so the null is rejected.

SPSS Cross-Tabulation Procedure

Exhibit 14-9
In another type of chi-square, the 2 x 2 table, a correction factor known as Yates’ correction for continuity is
applied when sample sizes are greater than 40 or when the sample is between 20 and 40 and the values of Ei
are 5 or more.
When the continuity correction is applied to the data shown in Exhibit 14-9, a chi-square of 5.25 is obtained.
The observed level of significance for this value is .02192. If the level of significance were set at .01, we would
accept the null hypothesis. However, had we calculated chi-square without the correction, the value would
have been 6.25 with an observed level of significance of .01242. The literature is in conflict over the merits of
Yates’ correction.
The Mantel-Haenszel test and the likelihood ratio also appear in Exhibit 14-9. The former is used with ordinal
data; the latter, based on maximum likelihood theory, produces results similar to Pearson’s chi-square.
Two-Related-Samples Tests
The two-related samples tests concern those situations in which persons, objects, or events are
closely matched or the phenomena are measured twice. For instance, one might compare the
consumption of husbands and wives.
Both parametric and nonparametric tests are applicable under these conditions
-Nonparametric
-Parametric
The two-related samples tests concern those situations in which persons, objects, or events are closely
matched or the phenomena are measured twice. For instance, one might compare the consumption of
husbands and wives.
Both parametric and nonparametric tests are applicable under these conditions.
Parametric
The t-test for independent samples is inappropriate here because of its assumption that observations are
independent. The problem is solved by a formula where the difference is found between each matched pair of
observations, thereby reducing the two samples to the equivalent of a one-sample case. In other words, there
are now several differences, each independent of the other, for which one can compute various statistics.
Nonparametric Tests
The McNemar test may be used with either nominal or ordinal data and is especially useful with before-after
measurement of the same subjects.

Sales Data for Paired-Samples t-Test

Exhibit 14-10 shows two years of Forbes sales data (in millions of dollars) from 10 companies. The next slide
illustrates the hypothesis test.
Paired-Samples t-Test Example

The null hypothesis states that there is no difference in sales data between years one and two. The alternate
hypothesis states that there is a difference.
The matched or paired-sample t-test is chosen because there are repeated measures on each company, the
data are not independent, and the measurement is ratio.
Enter Appendix Exhibit C-2 with d.f. = 9, two-tailed test, alpha = .01, and find the critical value of 3.25.
The calculated value is greater than the critical value so the null is rejected. We conclude that there is a
significant difference between the two years of sales.
SPSS Output for Paired-Samples t-Test

A computer solution to the problem is illustrated in Exhibit 14-11. Notice that an observed significance level
is printed for the calculated t value (highlighted). The observed significance level is the probability value
compared to the significance level chosen for testing and on this basis, the null hypothesis is either rejected
or not rejected.

Related Samples Nonparametric Tests: McNemar Test

The McNemar test may be used with either nominal or ordinal data and is especially useful with before-after
measurement of the same subjects. One can test the significance of any observed change by setting up a
fourfold table of frequencies to represent the first and second set of responses. An example is provided in the
slide.
Since A + D represents the total number of people who changed (B and C are no change responses), the null
hypothesis is that ½ (A+D) cases change in one direction and the same proportion in the other direction. The
McNemar test uses a transformation of the chi-square test.
Use formula from page 515
The minus 1 in the equation is a correction for continuity since the chi-square is a continuous distribution and
the observed frequencies represent a discrete distribution

Related Samples Nonparametric Tests: McNemar Test

The McNemar test may be used with either nominal or ordinal data and is especially useful with before-after
measurement of the same subjects. One can test the significance of any observed change by setting up a
fourfold table of frequencies to represent the first and second set of responses. An example is provided in the
slide.
Since A + D represents the total number of people who changed (B and C are no change responses), the null
hypothesis is that ½ (A+D) cases change in one direction and the same proportion in the other direction. The
McNemar test uses a transformation of the chi-square test.
Use formula from page 516
The minus 1 in the equation is a correction for continuity since the chi-square is a continuous distribution and
the observed frequencies represent a discrete distribution.

k-Independent-Samples Tests: ANOVA


Tests the null hypothesis that the means of three or more populations are equal.
One-way: Uses a single-factor, fixed-effects model to compare the effects of a treatment or factor
on a continuous dependent variable.
In a fixed-effects model, the levels of the factor are established in advance and the results are not
generalizable to other levels of treatment.
To use ANOVA, certain conditions must be met.

• The samples must be randomly selected from normal populations and the populations
should have equal variances.

• The distance from one value to its group’s mean should be independent of the distances of
other values to that mean.

Unlike the t-test, which uses sample standard deviations, ANOVA uses squared deviations of the variance to
that computation of distances of the individual data points from their own mean or from the grand mean can
be summed (recall that standard deviations sum zero).
In an ANOVA model, each group has its own mean and values that deviate from that mean. The total deviation
is the sum of the squared differences between each data point and the overall grand mean.
The total deviation of any particular data point may be partitioned into between-groups variance and within-
groups variance. The between-groups variance represents the effect of the treatment or factor. The
differences of between-groups means imply that each group was treated differently and the treatment will
appear as deviations of the sample means from the grand mean.
The within-groups variance describes the deviations of the data points within each group from the sample
mean. It is often called error.

ANOVA Example

Exhibit 14-13, top two parts


Malaysia Airlines
The test statistic for ANOVA is the F ratio.
Use formula from page 497
To compute the F ratio, the sum of the squared deviations for the numerator and denominator are divided by
their respective degrees of freedom. By dividing, we are computing the variance as an average or mean, thus
the term mean square. The degrees of freedom for the numerator, the mean square between groups, are one
less than the number of group (k-1). The degrees of freedom for the denominator, the mean square within
groups, are the total number of observations minus the number of groups (n-k).
If the null is true, there should be no difference between the population means and the ratio should be close
to 1. If the population means are not equal, the F should be greater than 1. The F distribution determines the
size of the ratio necessary to reject the null for a particular sample size and level of significance.

ANOVA Example Continued

To illustrate, consider the report about the quality of in-flight service on various carriers from the US to
Europe. Three airlines are compared. The data are shown in Exhibit 14-12. The dependent variable is service
rating and the factor is airline.
he null hypothesis states that there is no difference in the service rating score between airlines.
ANOVA and the F test is chosen because we have k independent samples, can accept the assumptions of
analysis of variance, and have interval data for the dependent variable. The significance level is .05. The
calculated F value is 28.30 (see summary table in the last slide).
Enter Appendix Exhibit C-9 with d.f. = 2, 57, and find the critical value of 3.16.
The calculated value is greater than the critical value so the null is rejected. We conclude that there is a
significant difference in flight service ratings.
Note that the p value provided in the summary table can also be used to reject the null.
Post Hoc: Scheffe’s S Multiple Comparison Procedure

Exhibit 14-13, bottom part


With an ANOVA, we cannot tell which pairs are not equal. We can use a post hoc test to determine where the
differences lie.
These tests find homogeneous subsets of means that are not different from each other. Multiple comparison
tests use group means and incorporate the MS error term of the F ratio. Together they produce confidence
intervals for the population means and a criterion score. Differences between the mean values may be
compared.
There are more than a dozen such tests. The exhibit in the slide uses Scheffe’s test. It is a conservative test
that is robust to violations of assumptions. In the example, all the differences between the pairs of means
exceed the critical difference criterion.
Multiple Comparison Procedures

Exhibit 14-14 There are several multiple comparison procedures one can use after rejecting the null with the F ratio. This
slide compares the available tests.

ANOVA Plots

Exhibit 14-15
In this exhibit, plots illustrate the comparison. The means plot shows relative differences among the three
levels of the factor. The means by standard deviations plot reveals lower variability in the opinions recorded
by the hypothetical Lufthansa and Cathay Pacific passengers. These two groups are sharply divided on the
quality of in-flight service and that is apparent in the plot.

Two-Way ANOVA Example

A two-way ANOVA is used to estimate how the mean of a quantitative variable changes
according to the levels of two categorical variables. Use a two-way ANOVA when you want to
know how two independent variables, in combination, affect a dependent variable
you could use a two-way ANOVA to understand whether there is an interaction between gender
and educational level on test anxiety amongst university students, where gender
(males/females) and education level (undergraduate/postgraduate) are your independent variables,
and test anxiety is your dependent
All data are hypothetical
Exhibit 14-16
Recall that in Exhibit 14-12 data were entered for the variable seat selection: economy and business-class
travelers. If we add this factor to our model, we have a two-way analysis of variance. We can now answer
three questions:

• Are differences in flight service ratings attributable to airlines?

• Are differences in flight service ratings attributable to seat selection?

• Do the airline and seat selections interact with respect to flight service ratings?

Exhibit 14-16, shown in the slide, tests the hypotheses for these questions. The significance level chosen is
.01. First, we consider the interaction effect of airline by seat selection. The null is accepted (p value is not
significant). But note that there are significant differences by airline and by seat selection.
When k independent samples are collected with nominal data, the chi-square is the appropriate
nonparametric technique. The Kruskal-Wallis test is appropriate for ordinal scale data or interval data tat do
not meet the F-test assumptions.

Two-way Analysis of Variance Plots

k-Related-Samples Tests
More than two levels in grouping factor
Observations are matched
Data are interval or ratio

The k-sample Anderson-Darling test is a nonparametric statistical procedure that tests the
hypothesis that the populations from which two or more groups of data were drawn are identical.
Each group should be an independent random sample from a population.
In test marketing experiments or ex post facto designs with k samples, it is often necessary to measure
subjects several times. These repeated measures are called trials.
The repeated-measures ANOVA is a special type of n-way analysis of variance. In this design, the repeated
measures of each subject are related just as they are in the related t-test where only two measures are
present. In this sense, each subject serves as its own control requiring a within-subjects variance effect to be
assessed differently than the between-groups variance in a factor like airline or seat selection.
This model is presented in Exhibit 14-18.
Summary Tables for Repeated-Measures ANOVA

Exhibit 14-18

Null hypotheses:

1) Airline: A1 = A2 = A3

2) Ratings: R1 = R2

3) Rating * Airline: (R2A1 - R2A2 - R2A3) = ((R1A1 - R1A2 - R1A3)

The F test for repeated measures is chosen because we have related trials on the dependent variable for k
samples, accept the assumptions of analysis of variance, and have interval data.

The significance level is .05.

The calculated values are shown in the exhibit.

The critical test values come from Appendix Exhibit C-9; they are 3.16 and 4.01.

Based on the results, all three null hypotheses are rejected. There are significant differences in all three
cases.
Repeated Measures ANOVA Plot

You might also like