0% found this document useful (0 votes)
2 views62 pages

Chapter 12

The document discusses statistical methods for comparing multiple proportions, including tests of independence and goodness of fit. It presents various examples with observed and expected frequencies, chi-square calculations, and conclusions based on p-values. The results indicate whether population proportions are equal or differ significantly across different groups.

Uploaded by

smadav1969
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views62 pages

Chapter 12

The document discusses statistical methods for comparing multiple proportions, including tests of independence and goodness of fit. It presents various examples with observed and expected frequencies, chi-square calculations, and conclusions based on p-values. The results indicate whether population proportions are equal or differ significantly across different groups.

Uploaded by

smadav1969
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Chapter 12

Comparing Multiple Proportions, Test of Independence and Goodness of Fit

Solutions

1. H0:

Ha: Not all population proportions are equal

Observed Frequencies (fij)

Tota
1 2 3 l

15
Yes 150 0 96 396

15 10
No 100 0 4 354

30 20
Total 250 0 0 750

Expected Frequencies (eij)

Tota
1 2 3 l

132. 158. 105.


Yes 0 4 6 396

118. 141.
No 0 6 94.4 354

Total 250 300 200 750

Chi-Square Calculations (fij – eij)2 / eij

T
ot
1 2 3 al

. . 3.
Ye 2. 4 8 7
s 45 5 7 7

No 2. . . 4.
5 9 2
75 0 8 2

Degrees of freedom = k – 1 = (3 – 1) = 2.

Using the table with df = 2, = 7.99 shows the p-value is


between .025 and .01.

Using software, the p-value corresponding to = 7.99 is .0184.

p-value .05; reject H0. Conclude not all population proportions are
equal.

2. a.

b. Multiple comparisons

For 1 versus 2

df = k –1 = 3 – 1 = z = 5.991

Differenc Critical Significant Diff >


Comparison pi pj e ni nj Value CV

1 vs. 2 .60 .50 .10 250 300 .1037

1 vs. 3 .60 .48 .12 250 200 .1150 Yes

2 vs. 3 .50 .48 .02 300 200 .1117

Only one comparison is significant, 1 versus 3. The others are not significant.
We can conclude that the population proportions differ for populations 1 and
3.

3. a. H0:
Ha: Not all population proportions are equal

b. Observed Frequencies (fij)


Flight Delta United US Airways Total

Delayed 39 51 56 146

On time 261 249 344 854

Total 300 300 400 1000

Expected Frequencies (eij)

Fli D U US T
gh e ni Air o
t lt te w t
a d ay a
s l

De 1
lay 4 4 58 4
ed 3 3. .4 6
. 8
8

On 2 2 34 8
ti 5 5 1. 5
m 6 6. 6 4
e . 2
2

Tot 3 3 40 1
al 0 0 0 0
0 0 0
0

Chi-Square Calculations (fij – eij)2 / eij

Flight Delt Unite US Tota


a d Airwa l
ys

Delaye .53 1.18 .10 1.8


d 1

On .09 .20 .02 .31


time
Degrees of freedom = k – 1 = (3 – 1) = 2.

Using the table with df = 2, = 2.12 shows the p-value is greater


than .10.

Using software, the p-value corresponding to = 2.12 is .3465.

p-value > .05; do not reject H0. We are unable to reject the null
hypothesis that the population proportions are the same.

c.

Overall

4. a. H0:
Ha: Not all population proportions are equal

b. Observed Frequencies (fij)

Component A B C Total

Defective 15 20 40 75

Good 485 480 460 1425

Total 500 500 500 1500

Expected Frequencies (eij)

Compone Tota
nt A B C l

Defective 25 25 25 75

47 47 47 142
Good 5 5 5 5

50 50 50 150
Total 0 0 0 0

Chi-Square Calculations (fij – eij)2 / eij

Co A B C T
m o
po t
ne a
nt l

1
De 4 1 9 4
fec . . . .
tiv 0 0 0 0
e 0 0 0 0

0
. . . .
Go 2 0 4 7
od 1 5 7 4

Degrees of freedom = k – 1 = (3 – 1) = 2

Using the table with df = 2, = 14.74 shows the p-value is less


than .005.

Using software, the p-value corresponding to = 14.74 is .0006.

Because the p-value < .05; reject H0. Conclude that the three
suppliers do not provide equal proportions of defective components.

c.

Multiple comparisons

For Supplier A Versus Supplier B

df = k –1 = 3 – 1 = z = 5.991

Significant Diff >


Comparison pi pj Difference ni nj Critical Value CV
.0 50 50
A vs. B 3 .04 .01 0 0 .0284

.0 50 50
A vs. C 3 .08 .05 0 0 .0351 Yes

.0 50 50
B vs. C 4 .08 .04 0 0 .0366 Yes

Supplier A and supplier B are both significantly different from


supplier C. Supplier C can be eliminated on the basis of a
significantly higher proportion of defective components. Since
suppliers A and supplier B are not significantly different in terms of
the proportion defective components, both suppliers should remain
candidates for use by Benson.

5. H0: p1= p2= p 3


Ha: Not all population proportions are equal

Observed Frequencies (fij)

Carnegie Classification

Type of Moderate Higher Highest Tota


University Research Research Research l
Activity Activity Activity

Public 38 76 81 195

Not-for-profit 58 31 34 123
private

Total 96 107 115 318

Expected Frequencies (eij)

Carnegie Classification

Type of Moderate Higher Highest Tota


University Research Research Research l
Activity Activity Activity

Public 58.87 65.61 70.52 195


Not-for-profit 37.13 41.39 44.48 123
private

Total 96 107 115 318


Chi-Square Calculations (fij – eij)2 / eij

Carnegie Classification

Type of Moderate Higher Highest Total


University Research Research Research
Activity Activity Activity

Public 7.40 1.64 1.56 10.6


0

Not-for-profit 11.73 2.61 2.47 16.8


private 0

Total 19.13 4.25 4.03 27.4


0

= 27.40

Degrees of freedom = (r – 1)(c – 1) = (2 – 1)(3 – 1) = 2

Using the table with df = 2, = 27.40 shows the p-value is less


than .005.

Using software, the p-value corresponding to = 27.40 is .000001.

p-value < .05; reject H0. The proportion of public universities is not
equal in each Carnegie category. The largest differences between
actual and expected frequencies are in the moderate research
activity classification, for which the number of public schools is much
less than expected and the number of not-for-profit private schools is
much greater than expected.

6. a. 14% error rate


9% error rate

b. H0:
Ha:

Observed Frequencies (fij)

Retur Offic Offic Tota


n e1 e2 l

Error 35 27 62
Correc
t 215 273 488

250 300 550

Expected Frequencies (eij)

Retur Office Office Tota


n 1 2 l

Error 28.18 33.82 62

Correc 221.8 266.1


t 2 8 488

250 300 550

Chi-Square Calculations (fij – eij)2 / eij

Retur Offic Offic Tota


n e1 e2 l

3.0
Error 1.65 1.37 2

Correc
t .21 .17 .38

df = k – 1 = (2 – 1) = 1.

Using the table with df = 1, = 3.41 shows the p-value is


between .10 and .05.

Using software, the p-value corresponding to = 3.41 is .0648.

p-value < .10; reject H0. Conclude that the two offices do not have
the same population proportion error rates.

c. With two populations, a chi-square test for equal population


proportions has one degree of freedom. In this case, the test statistic
is always equal to z2. This relationship between the two test
statistics always provides the same p-value and the same conclusion
when the null hypothesis involves equal population proportions.
However, the use of the z-test statistic provides options for one–
tailed hypothesis tests about two population proportions while the
chi-square test is limited to two–tailed hypothesis tests about the
equality of the two population proportions.

7. a. H0:
Ha: Not all population proportions are equal

Observed Frequencies (fij)

Social Media United China Russia USA Total


Kingdom

Yes 480 215 343 640 1,678

No 320 285 357 360 1,322

800 500 700 1,000 3,000

Expected Frequencies (eij)

Social Media United China Russia USA Total


Kingdom

Yes 447.47 279.6 391.5 559.33 1,678


7 3

No 352.53 220.3 308.4 440.67 1,322


3 7

800 500 700 1,000 3,000

Chi-Square Calculations (fij – eij)2 / eij

Social Media United China Russia USA Total


Kingdom

Yes 2.36 14.95 6.02 11.63 34.96

No 3.00 18.98 7.63 14.77 44.38


2
χ =79.34
Degrees of freedom = df = k – 1 = (4 – 1) = 3.

Using the table with df = 3, = 79.34 shows the p-value is less


than .005.
Using software, the p-value corresponding to = 79.34 is essentially
0.

p-value .05; reject H0. Conclude the population proportions are not
all equal.

b. United Kingdom 480/800 = .60

China 215/500 = .43


Russia 343/700 = .49
United 640/1000 = .64 (Largest with 64%
States of adults)
c. Multiple pairwise comparisons

where df = k –1 = 4 – 1 = 3 and = 7.815

Comparison pi pj Differenc ni nj CVij Diff > CVij


e

UK vs. C 0.60 0.43 0.17 800 500 0.0786 Yes

UK vs. R 0.60 0.49 0.11 800 700 0.0717 Yes

UK vs US 0.60 0.64 0.04 800 1000 0.0644

C vs R 0.43 0.49 0.06 500 700 0.0814

C vs US 0.43 0.64 0.21 500 1000 0.0750 Yes

R vs US 0.49 0.64 0.15 700 1000 0.0678 Yes

Only two comparisons are not significant: the difference in the


proportion of adults that use social media in the United Kingdom and
the United States is not significantly different, nor is the difference in
the proportion of adults that use social media in China and Russia.
All other comparisons show a significant difference.

8. H0: The distribution of defects is the same for all suppliers


Ha: The distribution of defects is not the same all suppliers

Observed Frequencies (fij)


Part Tested A B C Total

Minor defect 15 13 21 49

Major defect 5 11 5 21

Good 130 126 124 380

Total 150 150 150 450


Expected Frequencies (eij)

Part tested A B C Total

Minor defect 16.33 16.33 16.33 49

Major defect 7.00 7.00 7.00 21

Good 126.67 126.67 126.67 380

Total 150 150 150 450

Chi-Square Calculations (fij – eij)2 / eij

Part tested A B C Total

Minor defect .11 .68 1.33 2.12

Major defect .57 2.29 .57 3.43

Good .09 .00 .06 .15

Degrees of freedom = (r – 1)(k – 1) = (3 – 1)(3 – 1) = 4

Using the  table with df = 4,


2
= 5.70 shows the p-value is greater
than .10

Using software, the p-value corresponding to = 5.70 is .2227

p-value > .05; do not reject H0. Conclude that we are unable to reject
the hypothesis that the population distribution of defects is the same
for all three suppliers. There is no evidence that quality of parts from
one suppliers is better than either of the others two suppliers.
9. H0: The column variable is independent of the row variable
Ha: The column variable is not independent of the row variable

Observed Frequencies (fij)

A B C Total

P 20 44 50 114

Q 30 26 30 86

Total 50 70 80 200

Expected Frequencies (eij)

A B C Total

P 28.5 39.9 45.6 114

Q 21.5 30.1 34.4 86

Total 50 70 80 200

Chi-Square Calculations (fij – eij)2 / eij

A B C Total

P 2.54 .42 .42 3.38

Q 3.36 .56 .56 4.48

= 7.86

Degrees of freedom = (2–1)(3–1) = 2.

Using the table with df = 2, = 7.86 shows the p-value is


between .01 and .025.

Using software, the p-value corresponding to = 7.86 is .0196.

p-value .05; reject H0. Conclude that there is an association between


the column variable and the row variable. The variables are not
independent.

10. H0: The column variable is independent of the row variable


Ha: The column variable is dependent on the row variable

Observed Frequencies (fij)

A B C Total
P 20 30 20 70

Q 30 60 25 115

R 10 15 30 55

Total 60 105 75 240

Expected Frequencies (eij)

A B C Total

P 17.50 30.63 21.88 70

Q 28.75 50.31 35.94 115

R 13.75 24.06 17.19 55

Total 60 105 75 240


Chi-Square Calculations (fij – eij)2 / eij

A B C Total

P .36 .01 .16 .53

Q .05 1.87 3.33 5.25

R 1.02 3.41 9.55 13.99

= 19.77

Degrees of freedom = (r – 1)(c – 1) = (3 – 1)(3– 1) = 4.

Using the table with df = 4, = 19.77 shows the p-value is less


than .005.

Using software, the p-value corresponding to = 19.77 is .0006.

p-value .05; reject H0. Conclude that the column variable is not
independent of the row variable.

11. a. H0: Type of ticket purchased is independent of the type of flight


Ha: Type of ticket purchased is not independent of the type of flight

Expected Frequencies

e11 = 35.59 e12 = 15.41


e21 = 150.73 e22 = 65.27
e31 = 455.68 e32 =
197.32
Ticket Flight Observed Expected Chi-Square
Frequency (fi) Frequency (ei) (fi – ei)2 / ei

First Domestic 29 35.59 1.22

International 22 15.41 2.82

Business Domestic 95 150.73 20.61

International121 65.27 47.59

Full fare Domestic 518 455.68 8.52

International135 197.32 19.68


Totals: 920
100.43

Degrees of freedom = (r – 1)(c – 1) = (3 – 1)(2 – 1) = 2.

Using the table with df = 2, = 100.43 shows the p-value is less


than .005.

Using software, the p-value corresponding to = 100.43 is .0000.

p-value .05; reject H0. Conclude that the type of ticket purchased is
not independent of the type of flight. We can expect the type of
ticket purchased to depend upon whether the flight is domestic or
international.

b. Column Percentages

Type of Flight

Type of Ticket Domestic International


First class 4.5% 7.9%
Business class 14.8% 43.5%
Economy class 80.7% 48.6%
A higher percentage of first class and business class tickets are
purchased for international flights compared to domestic flights.
Economy class tickets are purchased more for domestic flights. The
first class or business class tickets are purchased for more than 50%
of the international flights; 7.9% + 43.5% = 51.4%.

12. a. H0: Employment plan is independent of the type of company


Ha: Employment plan is not independent of the type of
company

Observed Frequency (fij)

Employment Plan Private Public Total

Add employees 37 32 69

No change 19 34 53

Lay off employees 16 42 58

Total 72 108 180


Expected Frequency (eij)

Employment plan Private Public Total

Add employees 27.6 41.4 69

No change 21.2 31.8 53

Lay off employees 23.2 34.8 58

Total 72.0 108.0 180

Chi-Square Calculations (fij – eij)2 / eij

Employment Plan Private Public Total

Add employees 3.20 2.13 5.34

No change 0.23 0.15 0.38

Lay off employees 2.23 1.49 3.72

= 9.44

Degrees of freedom = (r – 1)(c – 1) = (3 – 1)(2 – 1) = 2

Using the table with df = 2, = 9.44 shows the p-value is


between .01 and 0.005.

Using software, the p-value corresponding to = 9.44 is .0089.

Because the p-value .05; reject H0. Conclude the employment plan
is not independent of the type of company. Thus, we expect
employment plan to differ for private and public companies.

b. Column probabilities: For example, 37/72 = .5139

Employment plan Private Public

Add employees .5139 .2963

No change .2639 .3148

Lay off employees .2222 .3889

Employment opportunities look to be much better for private


companies with over 50% of private companies planning to add
employees (51.39%). Public companies have the greater proportions
of no change and lay off employees planned. 38.89% of public
companies are planning to lay off employees over the next 12
months. 69/180 = .3833, or 38.33% of the companies in the survey
are planning to hire and add employees during the next 12 months.

13. H0: Interest in leaving job for more money is independent of the
employee generation
Ha: Interest in leaving job for more money is not independent of the
employee generation

Observed Frequencies (fij)

Leave Job for More Money? Baby Generation Millenni Total


Boomer X al

Yes 129 152 164 445

No 207 183 171 561

Total 336 335 335 1,006

Expected Frequencies (eij)

Leave Job for More Baby Generation Millenial Tota


Money? Boomer X l

Yes 148.6 148.2 148.2 445

No 187.4 186.8 186.8 561

Total 336 335 335 225

Chi-Square Calculations (fij – eij)2 / eij

Leave Job for More Baby Generation Millenial Tot


Money? Boomer X al

Yes 2.59 .10 1.69 4.3


8

No 2.06 .08 1.34 3.4


7

= 7.85

Degrees of freedom = (r – 1)(c – 1)= (2 – 1)(3 – 1) = 2.

table with df = 2,  = 7.85 shows the p-value is


2
Using the
between .01 and .025.
Using software, the p-value corresponding to = 7.85 with df = 2
is .0197.

p-value .05; reject H0. Conclude interest in leaving job for more
money is not independent of the employee generation.

14. a. H0: Quality rating is independent of the education of the owner


Ha: Quality rating is not independent of the education of the owner

Observed Some HS Some College Total


Frequencies HS Grad College Grad
(fij)Quality Rating

Average 35 30 20 60 145

Outstanding 45 45 50 90 230

Exceptional 20 25 30 50 125

Total 100 100 100 200 500

Expected Frequencies (eij)

Quality Rating Some HS Some College Total


HS Grad College Grad

Average 29 29 29 58 145

Outstanding 46 46 46 92 230

Exceptional 25 25 25 50 125

Total 100 100 100 200 500

Chi-Square Calculations (fij – eij)2 / eij

Quality Rating Some HS HS Some College Total


Grad College Grad

Average 1.24 .03 2.79 .07 4.14

Outstanding .02 .02 .35 .04 .43

Exceptional 1.00 .00 1.00 .00 2.00

Degrees of freedom = (r – 1)(c – 1) = (3 – 1)(4 – 1) = 6.


Using the table with df = 6, = 6.57 shows the p-value is greater
than .10.

Using software, the p-value corresponding to = 6.57 is .3624.

p-value > .05; do not reject H0. We are unable to conclude that the
quality rating is not independent of the education of the owner. Thus,
quality ratings are not expected to differ with the education of the
owner.

b. Average: 145/500 = 29%

Outstanding 230/500 = 46%


Exceptional 125/500 = 25%
New owners look to be pretty satisfied with their new automobiles
with almost 50% rating the quality outstanding and over 70% rating
the quality outstanding or exceptional.

15. a. H0: Quality of management is independent of the reputation of


the company
Ha: Quality of management is not independent of the reputation of
the company
Observed Frequencies (fij)

Quality of Management Excellen Good Fair Total


t

Excellent 40 25 5 70

Good 35 35 10 80

Fair 25 10 15 50

Total 100 70 30 200

Expected Frequencies (eij)

Quality of Management Excellent Good Fair Total

Excellent 35.0 24.5 10.5 70

Good 40.0 28.0 12.0 80

Fair 25.0 17.5 7.5 50

Total 100 70 30 200

Chi-Square Calculations (fij – eij)2 / eij

Quality of Management Excellent Good Fair Total

Excellent .71 .01 2.88 3.61

Good .63 1.75 .33 2.71

Fair .00 3.21 7.50 10.71

Degrees of freedom = (r – 1)(c – 1) = (3 – 1)(3 – 1) = 4.

Using the table with df = 4, = 17.03 shows the p-value is less


than .005.

Using software, the p-value corresponding to = 17.03 is .0019.

p-value < .05; reject H0. Conclude that the rating for the quality of
management is not independent of the rating for the reputation of
the company.

b. Using the highest column probabilities, if the reputation of the


company is:
Excellent—There is a 40/100 = .40 chance the quality of
management will also be excellent.

Good—There is a 35/70 = .50 chance the quality of management will


also be good.

Fair—There is a 15/30 = .50 chance the quality of management will


also be fair.

The highest probabilities are that the two variables will have the
same ratings. Thus, the two ratings are associated.

16. a. Observed Frequency (fij)

Age of Respondent

Actress 18– 31– 45– Over


30 44 58 58 Total
s

Jessica Chastain 51 50 41 42 184

Jennifer 63 55 37 50 205
Lawrence

Emmanuelle 15 44 56 74 189
Riva

Quvenzhané 48 25 22 31 126
Wallis

Naomi Watts 36 65 62 33 196

Totals 213 239 218 230 900

The sample size is 900.

b. The sample proportion of movie fans who prefer each actress is:
The movie fans favored Jennifer Lawrence, but three other nominees
(Jessica Chastain, Emmanuelle Riva, and Naomi Watts) each were
favored by almost as many of the fans.

c. Expected Frequency (eij)

Age of Respondent

Actress 18–30 31– 45–58 Over Totals


44 58

Jessica Chastain 43.5 48.9 44.6 47.0 184

Jennifer Lawrence 48.5 54.4 49.7 52.4 205

Emmanuelle Riva 44.7 50.2 45.8 48.3 189

Quvenzhané 29.8 33.5 30.5 32.2 126


Wallis

Naomi Watts 46.4 52.0 47.5 50.1 196

Totals 213 239 218 230 900

Calculate for each cell in the table.

Age of Respondent

Actress 18–30 31– 45– More Than 58 Totals


44 58

Jessica Chastain 1.28 0.03 0.29 0.54 2.12

Jennifer 4.32 0.01 3.23 0.11 7.66


Lawrence

Emmanuelle 19.76 0.76 2.28 13.67 36.48


Riva

Quvenzhané 11.08 2.14 2.38 0.04 15.65


Wallis

Naomi Watts 2.33 3.22 4.44 5.83 15.82

Totals 38.77 6.16 12.6 20.20 77.74


1
With (5 – 1)(4 – 1) = 12 degrees of freedom, the p-value is
approximately 0.

p-value .05; reject H0. Attitude toward the actress who was most
deserving of the 2013 Academy Award for actress in a leading role is
not independent of age.

17. a. H0: Hours of sleep per night is independent of age


Ha: Hours of sleep per night is not independent of age

Observed Frequencies (fij)

Hours of Sleep 39 or 40 or Total


Younger Older

Fewer than 6 38 36 74

6 to 6.9 60 57 117

7 to 7.9 77 75 152

8 or more 65 92 157

Total 240 260 500

Expected Frequencies (eij)

Hours of Sleep 39 or 40 or Total


Younger Older

Fewer than 6 35.52 38.48 74

6 to 6.9 56.16 60.84 117

7 to 7.9 72.96 79.04 152

8 or more 75.36 81.64 157

Total 240 260 500

Chi-Square Calculations (fij – eij)2 / eij

Hours of Sleep 39 or Younger 40 or Older Total

Fewer than 6 .17 .16 .33

6 to 6.9 .26 .24 .50


7 to 7.9 .22 .21 .43

8 or more 1.42 1.31 2.74

= 4.01

Degrees of freedom = (r – 1)(c – 1) = (4 – 1)(2 – 1) = 3

Using the table with df = 3, = 4.01 shows the p-value is greater


than .10.

Using software, the p-value corresponding to = 4.01 is .2604.

p-value > .05; do not reject H0. Cannot reject the assumption that
age and hours of sleep are independent.

b. Because age does not appear to have an association on hours of


sleep, use the overall row percentages.

Fewer than 6 74/500 = .14 14.8%


8
6 to 6.9 117/500 = .23 23.4%
4
7 to 7.9 152/500 = .30 30.4%
4
8 or more 157/500 = .31 31.4%
4
30.4% + 31.4% = 61.8% of individuals get seven or more hours of
sleep a night.

18. Expected frequencies:

e11 = 11.81 e12 = 8.44 e13 = 24.75


e21 = 8.40 e22 = 6.00 e23 = 17.60
e31 = 21.79 e32 = 15.56 e33 = 45.65

Observed Frequency Expected Chi-Square


Frequency
Host A Host B (fi) (ei) (fi – ei)2 / ei

Con Con 24 11.81 12.57

Con Mixed 8 8.44 .02

Con Pro 13 24.75 5.58

Mixed Con 8 8.40 .02

Mixed Mixed 13 6.00 8.17

Mixed Pro 11 17.60 2.48

Pro Con 10 21.79 6.38

Pro Mixed 9 15.56 2.77

Pro Pro 64 45.65 7.38

= 45.36

Degrees of freedom = (r – 1)(c – 1) = (3 – 1)(3 – 1) = 4

table with df = 2,  = 45.36 shows the p-value is less


2
Using the
than .005.

Using software, the p-value corresponding to = 45.36 is .0000.

p-value .01; reject H0. Conclude that the ratings of the two hosts are
not independent. The host responses are more similar than different
and they tend to agree or be close in their ratings.

19. a. Expected frequencies:

e1 = 200 (.40) = 80, e2 = 200 (.40) = 80, e3 = 200 (.20) = 40

Observed frequencies:

f1 = 60, f2 = 120, f3 = 20

k – 1 = 2 degrees of freedom
Using the table with df = 2, = 35 shows the p-value is less
than .005.

Using software, the p-value corresponding to = 35 is


approximately 0.

p-value .01; reject . Conclude the proportions differ from .40, .40,
and .20.

b. = 9.210

Reject H0 if 9.210

= 35; reject . Conclude the proportions differ from .40, .40,


and .20.

20. With n = 30 we will use six classes, each with the probability of .1667.

= 22.8 s = 6.27

The z values that create six intervals, each with probability .1667, are
–.97, –.43, 0, .43, .97.

z Cutoff Value of x

–.97 22.8 – .97 (6.27) = 16.74

–.43 22.8 – .43 (6.27) = 20.10

0 22.8 + 0 (6.27) = 22.80

.43 22.8 + .43 (6.27) = 25.50

.97 22.8 + .97 (6.27) = 28.86

Interval Observed Expected Difference


Frequency Frequency

Less than 16.74 3 5 –2

16.74–20.10 7 5 2

20.10–22.80 5 5 0

22.80–25.50 7 5 2

25.50–28.86 3 5 –2

28.86 and up 5 5 0
2 (−2 )2 ( 2 )2 ( 0 )2 ( 2 )2 (−2 )2 ( 0 )2 16
χ= + + + + + = =3.20
5 5 5 5 5 5 5
Degrees of freedom = k – p – 1 = 6 – 2 – 1 = 3

Using the table with df = 3, = 3.20 shows the p-value is greater


than .10.

Using software, the p-value corresponding to = 3.20 is .3618.

p-value > .05; do not reject . The claim that the data come from a
normal distribution cannot be rejected.

21. H0: pABC = .29, pCBS = .28, pNBC = .25, pIND = .18
Ha: The proportions are not pABC = .29, pCBS = .28, pNBC = .25, pIND
= .18

Expected frequencies:

300(.29) = 87, 300(.28) = 84

300(.25) = 75, 300(.18) = 54

e1 = 87, e2 = 84, e3 = 75, e4 = 54

Observed frequencies:

f1 = 95, f2 = 70, f3 = 89, f4 = 46

k – 1 = 3 degrees of freedom

Using the table with df = 3, = 6.87 shows the p-value is


between .05 and .10.

Using software, the p-value corresponding to = 6.87 is .0762.

p-value > .05; do not reject H0. There has not been a significant
change in the viewing audience proportions.

22.

Category Hypothesize Observed Expected Chi-Square


d Proportion Frequency (fi) Frequency (ei)(fi – ei)2 / ei

Blue .24 105 120 1.88

Brown .13 72 65 .75

Green .20 89 100 1.21

Orange .16 84 80 .20

Red .13 70 65 .38

Yellow .14 80 70 1.43

Total: 500 = 5.85

k – 1 = 6 – 1 = 5 degrees of freedom

Using the table with df = 5, = 5.85 shows the p-value is greater


than .10

Using software, the p-value corresponding to  = 5.85 is .3211


2

p-value > .05; do not reject H0. We cannot reject the hypothesis that
the overall percentages of colors in the population of M&M milk
chocolate candies are .24 blue, .13 brown, .20 green, .16 orange, .13
red and .14 yellow.

23. Expected frequencies:

20% each n = 60

e1 = 12, e2 = 12, e3 = 12, e4 = 12, e5 = 12

Observed frequencies:

f1 = 5, f2 = 8, f3 = 15, f4 = 20, f5 = 12

k – 1 = 4 degrees of freedom.

Using the table with df = 4, = 11.50 shows the p-value is


between .01 and .025.

Using software, the p-value corresponding to = 11.50 is .0215.


p-value < .05; reject . Conclude the largest companies differ in
performance from the 1000 companies. In general, the largest
companies did not do as well as others. 15 of 60 companies (25%) are
in the middle group and 20 of 60 companies (33%) are in the next
lower group. These both are greater than the 20% expected. Relative
few large companies are in the top A and B categories.

24. a. H0:
Ha: Not all proportions are equal

Observed Frequency (fi)

Sunday Monda Tuesda Wednesd Thursd Frida Saturd


y y ay ay y ay
66 50 53 47 55 69 80
Expected Frequency (ei) ei = 1/7(420) = 60

Sunday Monda Tuesda Wednesd Thursd Frida Saturd


y y ay ay y ay
60 60 60 60 60 60 60
Chi-Square Calculations (fi – ei)2 / ei

Sunday Monda Tuesda Wednesd Thursd Frida Saturd


y y ay ay y ay
.60 1.67 .82 2.82 .42 1.35 6.67

= 14.33

Degrees of freedom = (k – 1) = ( 7 – 1) = 6

Using the table with df = 6, = 14.33 shows the p-value is


between .05 and .025.

Using software, the p-value corresponding to = 14.33 is .0262

p-value .05; reject H0. Conclude the proportion of traffic accidents is


not the same for each day of the week.

b. Percentage of traffic accidents by day of the week

Sunday 66/420 15.71


= .1571 %
Monday 50/420 11.90
= .1190 %
Tuesday 53/420 12.62
= .1262 %
Wednesd 47/420 11.19
ay = .1119 %
Thursday 55/420 13.10
= .1310 %
Friday 69/420 16.43
= .1643 %
Saturday 80/420 19.05
= .1905 %
Saturday has the highest percentage of traffic accident (19%).
Saturday is typically the late night and more social day/evening of
the week. Alcohol, speeding and distractions are more likely to affect
driving on Saturdays. Friday is the second highest with 16.43%.

25. = 71 s = 17 n = 25

Use five classes.

Percenta z Data Value


ge

20.00% –.8 71–.84(17) =


4 56.72

40.00% –.2 71–.25(17) =


5 66.75

60.00% .25 71+.25(17) =


75.25

80.00% .84 71+.84(17) =


85.28

Interval Observed Expected


Frequency Frequency

Less than 56.72 7 5


56.72–66.75 7 5

66.75–75.25 1 5

75.25–85.28 1 5

85.28 and up 9 5

= 11.20

Degrees of freedom = k – p – 1 = 5 – 2 – 1 = 2.

Using the table with df = 2, = 11.20 shows the p-value is less


than .005.

Using software, the p-value corresponding to = 11.20 is .0037.

p-value .01; reject H0. Conclude the distribution does not have a
normal probability distribution.

26. = 24.5 s = 3 n = 30

Use six classes,

Percentage z Data Value

16.67% –.97 24.5–.97(3) =


21.59

33.33% –.43 24.5–.43(3) =


23.21

50.00% .00 24.5+.00(3) =


24.50

66.67% .43 24.5+.43(3) =


25.79

83.33% .97 24.5+.97(3) =


27.41

Interval Observed Expected


Frequency Frequency

Less than 21.59 5 5


21.59–23.21 4 5

23.21–24.50 3 5

24.50–25.79 7 5

25.7927.41 7 5

27.41 up 4 5

= 2.80

Degrees of freedom = (k – p – 1) = 6 – 2 – 1 = 3.

Using the table with df = 3, = 2.80 shows the p-value is greater


than .10.

Using software, the p-value corresponding to = 2.80 is .4235.

p-value > .10; do not reject H0. The assumption of a normal


distribution cannot be rejected.

27. a.

Washington, D.C. 8.8%; Bridgeport, CT 11.7%; San Jose, CA 9%,


Lexington Park, MD 8.5%

b. H0:
Ha: Not all population proportions are equal

Observed Frequencies (fij)

Millionaire Bridgeport San Jose, Washington, Lexington Total


, CT CA D.C. Park, MD

Yes 44 35 35 34 148

No 356 350 364 366 1,436


Total 400 385 399 400 1,584

Expected Frequencies (eij)

Millionaire Bridgepor San Jose, Washington, Lexington Total


t, CT CA D.C. Park, MD

Yes 37.37 35.97 37.28 37.37 148

No 362.63 349.03 361.72 362.63 1,436

Total 400 385 399 400 1,584

Chi-Square Calculations (fij – eij)2 / eij

Millionaire Bridgepor San Jose, Washingto Lexington Total


t, CT CA n, D.C. Park, MD

Yes 1.18 .03 .14 .30 1.65

No .12 .00 .01 .03 .16

= 1.81

Degrees of freedom = k – 1 = (4 – 1) = 3

Using the table with df = 3, = 1.81 shows the p-value is greater


than .10

Using software, the p-value corresponding to = 1.81 is .6117

p-value > .05; do not reject H0. Cannot conclude that there is a
difference among the population proportion of millionaires for these
four cities.

28. a. H0:
Ha: Not all population proportions are equal

Observed Frequencies (fij)

Quality First Second Third Total

Good 285 368 176 829

Defective 15 32 24 71

Total 300 400 200 900

Expected Frequencies (eij)


Quality First Second Third Total

Good 276.33 368.44 184.22 829

Defective 23.67 31.56 15.78 71

Total 300 400 200 900

Chi-Square Calculations (fij – eij)2 / eij

Quality First Second Third Total

Good .27 .00 .37 .64

Defective 3.17 .01 4.28 7.46

Degrees of freedom = k – 1 = (3 – 1) = 2.

Using the table with df = 2, = 8.10 shows the p-value is


between .025 and .01.

Using software, the p-value corresponding to = 8.10 is .0174.

p-value .05; reject H0. Conclude the population proportion of good


parts is not equal for all three shifts. The shifts differ in terms of
production quality.

b.

df = k –1 = 3 – 1 = 2

Comparison pi pj Differenc ni nj Critical Significant


e Value Diff > CV

1 vs. 2 .95 .92 .03 300 400 .0453

1 vs. 3 .95 .88 .07 300 200 .0641 Yes

2 vs. 3 .92 .88 .04 400 200 .0653


Shifts 1 and 3 differ significantly with shift 1 producing better quality
(95%) than shift 3 (88%). The study cannot identify shift 2 (92%) as
better or worse quality than the other two shifts. Shift 3, at 7% more
defectives than shift 1 should be studied to determine how to
improve its production quality.

29. Let p1 = population proportion of visitors who rate the Louvre Museum
as spectacular.

p2 = population proportion of visitors who rate the National Museum


in China as spectacular

p3 = population proportion of visitors who rate the Metropolitan


Museum of Art as spectacular

p4 = population proportion of visitors who rate the Vatican Museums


as spectacular

p5 = population proportion of visitors who rate the British Museums


as spectacular

a. Point estimates of the population proportion of visitors who rated


each of these museums as spectacular are:

= 113/150 = .7533 is the point estimate of the population


proportion of visitors who rated the Louvre Museum as spectacular.

= 88/132 = .6667 is the point estimate of the population


proportion of visitors who rated the National Museum in China as
spectacular.

= 94/140 = .6714 is the point estimate of the population


proportion of visitors who rated the Metropolitan Museum of Art as
spectacular.

= 98/170 = .5765 is the point estimate of the population


proportion of visitors who rated the Vatican Museums as spectacular.

= 96/160 = .6000 is the point estimate of the population


proportion of visitors who rated the British Museum as spectacular.

b. H0: p1 = p2 = p3 = p4 = p5
Ha: Not all population proportions are equal
Observed Frequency (fij)

Louvre National Metropolitan Vatican


Museum Museum in Museum of Art Museums
China

Rated spectacular 113 88 94 98

Did not rate spectacular 37 44 46 72

Totals 150 132 140 170

Expected Frequency (eij)

Louvre National Metropolitan Vatican


Museum Museum in Museum of Art Museum
China

Rated spectacular 97.54 85.84 91.04 110.55

Did not rate 52.46 46.16 48.96 59.45


spectacular

Totals 150 140 160 170


Chi-Square (fij – eij)2 / eij

Louvr Nation Metropoli Vatican British Tot


e al tan Museu Museu al
Museu Museu Museum ms m
m m in of Art
China

Rated spectacular 2.45 .05 .10 1.42 .62 4.6


5

Did not rate 4.56 .10 .18 2.65 1.16 8.6


spectacular 4

= 13.29

Degrees of freedom = k – 1 = 5 – 1 = 4.

table with df = 4,  = 13.29 shows the p-value is


2
Using the
between .005 and .01.

Using software, the p-value corresponding to = 13.29 is .00996.

p-value ≤ .05; reject H0. We conclude that the population proportion


of visitors who rated the museum as spectacular differs for these
five museums.
30. a. H0: The preferred pace of life is independent of gender
Ha: The preferred pace of life is not independent of gender

Observed Frequency (fij)

Gender

Preferred Pace of Life Male Female Total

Slower 230 218 448

No Preference 20 24 44

Faster 90 48 138

Total 340 290 630

Expected Frequency (eij)

Gender

Preferred Pace of Life Male Female Total

Slower 241.78 206.22 448

No Preference 23.75 20.25 44

Faster 74.48 63.52 138

Total 340 290 630


Chi-Square Calculations (fij – eij)2/ eij

Gender

Preferred Pace of Life Male Female Total

Slower .57 .67 1.25

No preference .59 .69 1.28

Faster 3.24 3.79 7.03

Degrees of freedom = (r – 1)(c – 1) = (3 – 1)(2 – 1) = 2.

Using the table with df = 2, = 9.56 shows the p-value is less


than .01.

Using software, the p-value corresponding to = 9.56 is .0084.

p-value < .05; reject H0. The preferred pace of life is not independent
of gender. Thus, we expect men and women differ with respect to
the preferred pace of life.

b. Percentage responses for each gender

Gender

Preferred Pace of Life Male Female

Slower 67.65 75.17

No preference 5.88 8.28

Faster 26.47 16.55

The highest percentages are for a slower pace of life by both men
and women. However, 75.17% of women prefer a slower pace
compared to 67.65% of men and 26.47% of men prefer a faster pace
compared to 16.55% of women. More women prefer a slower pace
while more men prefer a faster pace.

31. H0: Church attendance is independent of age


Ha: Church attendance is not independent on age

Observed Frequencies (fij)


Age

Church Attendance 20–29 30–39 40–49 50–59 Total

Yes 31 63 94 72 260

No 69 87 106 78 340

Total 100 150 200 150 600

Expected Frequencies (eij)

Age

Church Attendance 20–29 30–39 40–49 50–59 Total

Yes 43 65 87 65 260

No 57 85 113 85 340

Total 100 150 200 150 600

Chi-Square (fij – eij)2/ eij

Age

Church Attendance 20–29 30–39 40–49 50–59 Total

Yes 3.51 .06 .62 .75 4.94

No 2.68 .05 .47 .58 3.78

Degrees of freedom = (r – 1)(c – 1) = (2 – 1)(4 – 1) = 3

Using the table with df = 3, = 8.72 shows the p-value is


between .025 and .05.

Using software, the p-value corresponding to = 8.72 is .0333.

p-value .05; reject . Conclude church attendance is not


independent of age.

Church Attendance by Age Group

20–29
31/100 31
%
30–39
63/150 42
%
40–49
94/200 47
%
50–59
72/150 48
%
Church attendance increases as individuals grow older.

32. H0: The county with the emergency call is independent of the day of
week
Ha: The county with the emergency call is not independent of the day
of week

Observed Frequencies (fij)

Day of Week

County Sun Mon Tues Wed Thu Fri Sat Total

Urban 61 48 50 55 63 73 43 393

Rural 7 9 16 13 9 14 10 78

Total 68 57 66 68 72 87 53 471

Expected Frequencies (eij)

Day of Week

Count Sun Mon Tue Wed Thu Fri Sat Total


y

Urban 56.74 47.56 55.07 56.74 60.08 72.5 44.2 393


9 2

Rural 11.26 9.44 10.93 11.26 11.92 14.4 8.78 78


1

Total 68 57 66 68 72 87 53 471

Chi-Square (fij – eij)2/ eij

Day of Week

County Sun Mon Tue Wed Thu Fri Sat Total


Urban .32 .00 .47 .05 .14 .00 .03 1.02

Rural 1.61 .02 2.35 .27 .72 .01 .17 5.15

= 6.17

Degrees of freedom = (r – 1)(c – 1) = (2 – 1)(7 – 1) =6.

Using the table with df = 6, = 6.17 shows the p-value is greater


than .10.

Using software, the p-value corresponding to = 6.17 is .4044.

p-value > .05; do not reject H0. The assumption of independence


cannot be rejected. The county with the emergency call does not vary
or depend upon the day of the week.

33. a. The sample size is very large: 6448

b.
Observed Frequency (fij)

Country

Response Great France Italy Spain Germa United States


Britain ny

Strongly favor 141 161 298 133 128 204

Favor 348 366 309 222 272 326

Oppose 381 334 219 311 322 316

Strongly oppose 217 215 219 443 389 174

Total 1,087 1,076 1,045 1,109 1,111 1,020

Expected Frequency (eij)

Country

Response Great Britain France Italy Spain Germa United States


ny

Strongly favor 180 178 173 183 183 168

Favor 311 307 299 317 318 291

Oppose 317 315 305 324 324 298

Strongly oppose 279 276 268 285 286 263

Total 1,087 1,076 1,04 1,109 1,111 1,020


5

Chi-Square (fij – eij)2/ eij

Country

Response Great Britain Franc Italy Spai Germa United


e n ny States

Strongly favor 8.45 1.62 90.3 13.6 16.53 7.71


2 6

Favor 4.40 11.34 0.33 28.4 6.65 4.21


7

Oppose 12.92 1.15 24.2 0.52 0.01 1.09


5
Strongly oppose 13.78 13.48 8.96 87.5 37.09 30.12
9

= 424.65

Degrees of freedom = (r – 1)(c – 1) = (4 – 1)(6 – 1) = 15

The p-value is approximately 0.

p-value .05; reject H0. The attitude toward building new nuclear
power plants is not independent of the country. Attitudes can be
expected to vary with the country.

c. Use column percentages from the observed frequencies table to help


answer this question.

Country

Response Great Britain France Italy Spain Germa United


ny States

Strongly 13.0 15.0 28.5 12.0 11.5 20.0


favor

Favor 32.0 34.0 29.5 20.0 24.5 32.0

Oppose 35.0 31.0 21.0 28.0 29.0 31.0

Strongly 20.0 20.0 21.0 40.0 35.0 17.0


oppose

Total 100 100 100 100 100 100

Adding together the percentages of respondents who “Strongly


favor” and those who “Favor”, we find the following: Great Britain
45%, France 49%, Italy 58%, Spain 32%, Germany 36% and United
States 52%. Italy shows the most support for nuclear power plants
with 58% in favor. Spain shows the least support with only 32% in
favor. Only Italy and the United States show more than 50% of the
respondents in favor of building new nuclear power plants.
34.

Expected Frequencies for n = 344

Professional football e1 = .33*344 = 113.52


Baseball e2 = .15*344 = 51.60
Men’s college football e3 = .10*344 = 34.40
Auto racing e4 = .06*344 = 20.64
Men’s professional e5 = .05*344 = 17.20
basketball
Ice hockey e6 = .05*344 = 17.20
Other sports e7 = .26*344 = 89.44
Actual Frequencies

Professional football f1 = 111


Baseball f2 = 39
Men’s college football f3 = 46
Auto racing f4 = 14
Men’s professional f5 = 6
basketball
Ice hockey f6 = 20
Other sports f7 = 108

k – 1 = 6 degrees of freedom

Using the table with df = 6, = 20.78 shows the p-value is less


than .005.

Using Excel, the p-value corresponding to = 20.78 with df = 6


is .0020.

p-value < .05; reject H0. Yes, undergraduate students differ from the
general public with regard to their favorite sports. Men’s college
football is more popular among undergraduate students, and
baseball and auto racing are less popular among undergraduate
students. The difference between expected and actual number of
undergraduate students who responded other sports also suggests
that undergraduate students are interested in a broader range of
sports.
35. H0: The market shares for the seven small-car categories in Chicago
are .20, .17, .12, .10, .10, .08, .23
Ha: The market shares for the seven small-car categories in Chicago differ
from the above shares

Compact Car Hypothesized Observed Expected Chi-


Market Share Frequency Frequency Square (fi
– e i) 2 / e i

Honda Civic .20 98 80 4.05

Toyota Corolla .17 72 68 0.24

Nissan Sentra .12 54 48 0.75

Hyundai Elantra .10 44 40 0.40

Chevrolet Cruze .10 42 40 0.10

Ford Focus .08 25 32 1.53

Other .23 65 92 7.92

c2 = 14.99

Degrees of freedom = k – 1 = 7 – 1 = 6.

Using the table with df = 6, = 14.99 shows the p-value is


between .01 and .025.

Using software, the p-value corresponding to = 14.99 is .02.

p-value < .05; reject .

Conclude that the Chicago market shares for the seven categories
differ from the national market shares. In particular, the Chicago
market appears to have fewer purchases of the “Other” category
and more purchases of the Honda Civic.
36. = 76.83 s = 12.43

Interval Observed Expected


Frequency Frequency

Less than 62.54 5 5

62.54–68.50 3 5

68.50–72.85 6 5

72.85–76.83 5 5

76.83–80.81 5 5

80.81–85.16 7 5

85.16–91.12 4 5

91.12 up 5 5

=2

Degrees of freedom = k – p – 1 = 8 – 2 – 1 = 5.

Using the table with df = 5, = 2.00 shows the p-value is greater


than .10.

Using software, the p-value corresponding to = 2.00 is .8491.

p-value > .05; do not reject . The assumption of a normal


distribution cannot be rejected.

37. a.

x Observed Binomial Probability Expected


Frequencies n = 4, p = .30 Frequencies

0 30 .2401 24.01

1 32 .4116 41.16

2 25 .2646 26.46

3 10 .0756 7.56

4 3 .0081 .81

100 100.00

The expected frequency of x = 4 is .81. Combine x = 3 and x = 4


into one category so that all expected frequencies are 5 or more.

x Observed Frequencies Expected Frequencies


0 30 24.01

1 32 41.16

2 25 26.46

3 or 4 13 8.37

100 100.00

b. = 6.17

Degrees of freedom = k – 1 = 4 – 1 = 3.

Using the table with df = 3, = 6.17 shows the p-value is


greater than .10.

Using software, the p-value corresponding to = 6.17 is .1036.

p-value > .05; do not reject H0. Conclude that the assumption of a
binomial distribution cannot be rejected.
Case Solutions

Case Problem 1 A Bipartisan Agenda for Change

1. Descriptive statistics

Question: Should legislative pay be cut for every day the state budget is
late?

Yes No Totals

Democrat 22 14 36
Independent 10 9 19
Republican 39 6 45
Totals 71 29 100
Percentage responding yes: Democrat, 61.1%; Independent, 52.6%;
Republican, 86.7%

Preliminary conclusion: Because the percentage of Republicans who


answered yes is much greater than the percentage of Democrats or
Independents who answered yes, the classifications do not appear to be
independent.

Question: Should there be restrictions on lobbyists?

Yes No Totals

Democrat 21 15 36

Independent 15 4 19

Republican 34 11 45

Totals 70 30 100

Percentage responding yes: Democrat, 58.3%; Independent, 78.9%;


Republican, 75.6%

Preliminary conclusion: Because the percentage of Democrats who


answered yes is much less than the percentage of Independents or
Republicans who answered yes, the classifications do not appear to be
independent.

Question: Should there be term limits requiring legislators to serve no


more than a fixed number of years?

Yes No Totals

Democrat 17 19 36
Independent 10 9 19

Republican 32 13 45

Totals 59 41 100

Percentage responding yes: Democrat, 47.2%; Independent, 52.6%;


Republican, 71.1%

Preliminary conclusion: Because the percentage of Republicans who


answered yes is much higher than the percentage of Democrats or
Independents who answered yes, the classifications do not appear to be
independent.

2. The p-value associated with the test for independence is .006.


Because the p-value is less than the level of significance, .05, we can
reject the null hypothesis of independence.

3. The p-value associated with the test for independence is .156.


Because the p-value is greater than the level of significance, .05, we
cannot reject the null hypothesis of independence. Note that our
preliminary conclusion based on the descriptive statistics was incorrect in
this case.

4. The p-value associated with the test for independence is .078.


Because the p-value is greater than the level of significance, .05, we
cannot reject the null hypothesis of independence. Note that our
preliminary conclusion based on the descriptive statistics was incorrect in
this case.

5. There does appear to be broad support across all political lines with
regard to restrictions on lobbyists and requiring legislators to serve a fixed
maximum number of years. However, with regard to the issue of a
legislative pay cut, the response (yes and no) and party affiliation are not
independent; for this issue, the Republican respondents have a much
different perspective than Democrats or Independents

Case Problem 2 Fuentes Salty Snacks, Inc.

1. Let pi be the proportion of stores in sales region i that currently


carries Fuentes’ Candied Bacon Potato Chips where

i = 1 for New England


i = 2 for Mid-Atlantic
i = 3 for Midwest
i = 4 for Great Plains
i = 5 for South Atlantic
i = 6 for Deep South
i = 7 for Mountain
i = 8 for Pacific

The Great Plains, South Atlantic, and Deep South sales regions have the
highest penetration and the Pacific sales region has the lowest
penetration.

2. H0:

Ha: Not all population proportions are equal

Observed Frequencies (fij)

Carries Fuentes’ Candied Bacon Potato


Chips?

Does Not
Region Carries Carry Total

New England 27 13 40

Mid-Atlantic 29 11 40

Midwest 28 12 40

Great Plains 33 7 40

South Atlantic 32 8 40

Deep South 33 7 40

Mountain 25 15 40

Pacific 17 23 40

Total 224 96 320

Expected Frequencies (eij)

Carries Fuentes’ Candied Bacon Potato Chips?

Region Carries Does Not Carry Total

New 28 12 40
England

Mid-Atlantic 28 12 40

Midwest 28 12 40
Great Plains 28 12 40

South 28 12 40
Atlantic

Deep South 28 12 40

Mountain 28 12 40

Pacific 28 12 40

Total 224 96 320

Chi Square Calculations (fij – eij)2 / eij

Carries Fuentes’ Candied Bacon Potato


Chips?

Region Carries Does Not Carry Total

New England 0.0357 0.0833 0.1190

Mid-Atlantic 0.0357 0.0833 0.1190

Midwest 0.0000 0.0000 0.0000

Great Plains 0.8929 2.0833 2.9762

South Atlantic 0.5714 1.3333 1.9048

Deep South 0.8929 2.0833 2.9762

Mountain 0.3214 0.7500 1.0714

Pacific 4.3214 10.0833 14.4048

Total 7.0714 16.5000 23.5714

Degrees of freedom = k – 1 = (8 – 1) = 7

Using the table with df = 7, = 23.5714 shows the p-value is less


than .005

Using Excel, the p-value corresponding to = 23.5714 with df = 7


is .0014.

p-value .05, reject H0. Conclude test the hypothesis that the proportion of
grocery stores that currently carry Fuentes’ Candied Bacon Potato Chips is
equal across its eight U.S. sales regions.
3. The critical value for the Marascuilo pairwise comparison procedure
at a = .05 for each pair of Fuentes’ sales regions are provided below.

CV12 = 0.3838 CV13 = CV14 = CV15 =


0.3886 0.3577 0.3653
CV16 = 0.3577 CV17 = CV18 = CV23 =
0.3995 0.4038 0.3794
CV24 = 0.3477 CV25 = CV26 = CV27 =
0.3555 0.3477 0.3906
CV28 = 0.3950 CV34 = CV35 = CV36 =
0.3530 0.3607 0.3530
CV37 = 0.3953 CV38 = CV45 = CV46 =
0.3997 0.3272 0.3187
CV47 = 0.3650 CV48 = CV56 = CV57 =
0.3697 0.3272 0.3724
CV58 = 0.3771 CV67 = CV68 = CV78 =
0.3650 0.3697 0.4103
The absolute value for the pairwise difference in sample proportions for
each pair of Fuentes’ sales regions follows.

The Great Plains and Pacific sales regions are significantly different at a
= .05, and the Deep South and Pacific sales regions are significantly
different at a = .05. We conclude that the Great Plains and Deep South
sales regions have higher penetration of Fuentes’ Candied Bacon Potato
Chips than does the Pacific sales region.
Case Problem 3 Fresno Board Games

1. A bar chart of the outcomes for the first randomly selected die
shows some variation in frequency of occurrence across the six outcomes,
as do the sample proportions for outcomes for the first randomly selected

Die 1
120
100
Frequency

80
60
40
20
0
1 2 3 4 5 6

Outcome

die:

A bar chart of the outcomes for the second randomly selected die shows
little variation in frequency of occurrence across the six outcomes as do
the sample proportions for outcomes for the second randomly selected
die:

Die 2
120

100
Frequency

80

60

40

20

0
1 2 3 4 5 6

Outcome
A bar chart of the outcomes for the third randomly selected die shows
some variation in frequency of occurrence across the six outcomes as do
the sample proportions for outcomes for the third randomly selected die:

A bar chart of the outcomes for the fourth randomly selected die shows
little variation in frequency of occurrence across the six outcomes as do
the sample proportions for outcomes for the fourth randomly selected die:

Die 5
Die 3
Frequency

120
150
100 100
Frequency

80 50
60 0
1 2 3 4 5 6
40
Outcome
20
0
1 2 3 4 5 6

Outcome

A bar chart of the outcomes for the fifth randomly selected die shows little
variation in frequency of occurrence across the six outcomes as do the
sample proportions for outcomes for the fifth randomly selected die:

Die 4
120
Frequency

100
80
60
40
20
0
1 2 3 4 5 6
Outcome
Although the die each have some variation in frequency of occurrence of
the six possible outcomes, none of the die have any single outcome that
appears to be occurring much less frequently or much more frequently
than would be expected if the five dice are fair.

2. For the first randomly selected die:

H0: The population of outcomes has a multinomial distribution with


p1 = p2 = p3 = p4 = p5 = p6 = 1/6

Ha: The population of outcomes does not have a multinomial distribution


with
p1 = p2 = p3 = p4 = p5 = p6 = 1/6

Outcome Hypothesize Observed Expected Chi Square


d Proportion Frequency (fi) Frequency (ei) (fi – ei)2 / ei

1 .1667 101 83.33 3.7453

2 .1667 86 83.33 0.0853

3 .1667 73 83.33 1.2813

4 .1667 74 83.33 1.0453

5 .1667 75 83.33 0.8333

6 .1667 91 83.33 0.7053

Total: 500 = 7.696

k – 1 = 6 – 1 = 5 degrees of freedom

Using the table with df = 5, = 7.696 shows the p-value is greater than
.10.

Using Excel, the p-value corresponding to  = 7.696 with df = 5 is .1738.


2

p-value > .01, do not reject H0. We cannot reject the hypothesis that
outcomes for the first randomly selected die have a multinomial
distribution with p1 = p2 = p3 = p4 = p5 = p6 = 1/6. We cannot conclude
that the first randomly selected die is unfair.

For the second randomly selected die:

H0: The population of outcomes has a multinomial distribution with


p1 = p2 = p3 = p4 = p5 = p6 = 1/6
Ha: The population of outcomes does not have a multinomial distribution
with
p1 = p2 = p3 = p4 = p5 = p6 = 1/6

Outcome Hypothesized Observed Expected Chi Square


Proportion Frequency (fi) Frequency (ei) (fi – ei)2/ei

1 .1667 85 83.33 0.0333

2 .1667 93 83.33 1.1213

3 .1667 74 83.33 1.0453

4 .1667 80 83.33 0.1333

5 .1667 86 83.33 0.0853

6 .1667 82 83.33 0.0213

Total: 500 = 2.440

k – 1 = 6 – 1 = 5 degrees of freedom

Using the table with df = 5, = 2.440 shows the p-value is greater than
.10.

Using Excel, the p-value corresponding to  = 2.440 with df = 5 is .7855.


2

p-value > .01, do not reject H0. We cannot reject the hypothesis that
outcomes for the second randomly selected die have a multinomial
distribution with p1 = p2 = p3 = p4 = p5 = p6 = 1/6. We cannot conclude
that the second randomly selected die is unfair.

For the third randomly selected die:

H0: The population of outcomes has a multinomial distribution with


p1 = p2 = p3 = p4 = p5 = p6 = 1/6

Ha: The population of outcomes does not have a multinomial distribution


with
p1 = p2 = p3 = p4 = p5 = p6 = 1/6

Outcome Hypothesize Observed Expected Chi Square


d Proportion Frequency (fi) Frequency (ei) (fi – ei)2 / ei

1 .1667 96 83.33 1.9253

2 .1667 75 83.33 0.8333

3 .1667 85 83.33 0.0333

4 .1667 63 83.33 4.9613


5 .1667 87 83.33 0.1613

6 .1667 94 83.33 1.3653

Total: 500
= 9.280

k – 1 = 6 – 1 = 5 degrees of freedom

Using the table with df = 5, = 9.280 shows the p-value is greater than
.10.

Using Excel, the p-value corresponding to  = 9.280 with df = 5 is .0984.


2

p-value > .01, do not reject H0. We cannot reject the hypothesis that
outcomes for the third randomly selected die have a multinomial
distribution with p1 = p2 = p3 = p4 = p5 = p6 = 1/6. We cannot conclude
that the third randomly selected die is unfair.

For the fourth randomly selected die:

H0: The population of outcomes has a multinomial distribution with


p1 = p2 = p3 = p4 = p5 = p6 = 1/6

Ha: The population of outcomes does not have a multinomial distribution


with
p1 = p2 = p3 = p4 = p5 = p6 = 1/6

Outcome Hypothesize Observed Expected Chi Square


d Proportion Frequency (fi) Frequency (ei) (fi – ei)2 / ei

1 .1667 93 83.33 1.1213

2 .1667 98 83.33 2.5813

3 .1667 74 83.33 1.0453

4 .1667 78 83.33 0.3413

5 .1667 80 83.33 0.1333

6 .1667 77 83.33 0.4813

Total: 500
= 5.704

k – 1 = 6 – 1 = 5 degrees of freedom

Using the table with df = 5, = 5.704 shows the p-value is greater


than .10.
2
Using Excel, the p-value corresponding to  = 5.704 with df = 5 is .3361.
p-value > .01, do not reject H0. We cannot reject the hypothesis that
outcomes for the fourth randomly selected die have a multinomial
distribution with p1 = p2 = p3 = p4 = p5 = p6 = 1/6. We cannot conclude
that the fourth randomly selected die is unfair.

For the fifth randomly selected die:

H0: The population of outcomes has a multinomial distribution with


p1 = p2 = p3 = p4 = p5 = p6 = 1/6

Ha: The population of outcomes does not have a multinomial distribution


with
p1 = p2 = p3 = p4 = p5 = p6 = 1/6

Outcome Hypothesize Observed Expected Chi Square


d Proportion Frequency (fi) Frequency (ei) (fi – ei)2 / ei

1 .1667 73 83.33 1.2813

2 .1667 83 83.33 0.0013

3 .1667 94 83.33 1.3653

4 .1667 70 83.33 2.1333

5 .1667 87 83.33 0.1613

6 .1667 93 83.33 1.1213

Total: 500
= 6.064

k – 1 = 6 – 1 = 5 degrees of freedom

Using the table with df = 5, = 6.064 shows the p-value is greater


than .10.
2
Using Excel, the p-value corresponding to  = 6.064 with df = 5 is .3000.

p-value > .01, do not reject H0. We cannot reject the hypothesis that
outcomes for the fifth randomly selected die have a multinomial
distribution with p1 = p2 = p3 = p4 = p5 = p6 = 1/6. We cannot conclude
that the fifth randomly selected die is unfair.

The results of these hypothesis tests do not provide evidence that any
single outcome occurs much less frequently or much more frequently than
would be expected if the 5 six-sided dice are fair. Conclude that BBG is
producing fair dice.

You might also like